taisazero/socratic-debugging-benchmark
收藏资源简介:
--- license: mit task_categories: - text2text-generation - text-generation language: - en tags: - code pretty_name: Socratic Debugging Benchmark --- # Socratic Debugging Benchmark <p align="center"> <img src="img/socratic_debugging.png" height="400"> </p> The repository contains the dataset for the Socratic Debugging Benchmark accompanying the papers ["Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations"](https://aclanthology.org/2023.bea-1.57/) in proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Application at ACL 2023 and ["Can Language Models Employ the Socratic Method? Experiments with Code Debugging"](https://dl.acm.org/doi/10.1145/3626252.3630799) in the proceedings of SIGCSE'24. The dataset is also hosted on [Github](https://github.com/taisazero/socratic-debugging-benchmark). Please cite the following papers if you use this dataset: ``` @inproceedings{al-hossami-etal-2023-socratic, title = "Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations", author = "Al-Hossami, Erfan and Bunescu, Razvan and Teehan, Ryan and Powell, Laurel and Mahajan, Khyati and Dorodchi, Mohsen", booktitle = "Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023)", month = jul, year = "2023", address = "Toronto, Canada", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.bea-1.57", pages = "709--726", abstract = "Socratic questioning is a teaching strategy where the student is guided towards solving a problem on their own, instead of being given the solution directly. In this paper, we introduce a dataset of Socratic conversations where an instructor helps a novice programmer fix buggy solutions to simple computational problems. The dataset is then used for benchmarking the Socratic debugging abilities of GPT-based language models. While GPT-4 is observed to perform much better than GPT-3.5, its precision, and recall still fall short of human expert abilities, motivating further work in this area.", } ``` ``` @inproceedings{al-hossami-etal-2024-can, author = {Al-Hossami, Erfan and Bunescu, Razvan and Smith, Justin and Teehan, Ryan}, title = {Can Language Models Employ the Socratic Method? Experiments with Code Debugging}, year = {2024}, isbn = {9798400704239}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3626252.3630799}, doi = {10.1145/3626252.3630799}, abstract = {When employing the Socratic method of teaching, instructors guide students toward solving a problem on their own rather than providing the solution directly. While this strategy can substantially improve learning outcomes, it is usually time-consuming and cognitively demanding. Automated Socratic conversational agents can augment human instruction and provide the necessary scale, however their development is hampered by the lack of suitable data for training and evaluation. In this paper, we introduce a manually created dataset of multi-turn Socratic advice that is aimed at helping a novice programmer fix buggy solutions to simple computational problems. The dataset is then used for benchmarking the Socratic debugging abilities of a number of language models, ranging from fine-tuning the instruction-based text-to-text transformer Flan-T5 to zero-shot and chain of thought prompting of the much larger GPT-4. The code and datasets are made freely available for research at the link below.}, booktitle = {Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1}, pages = {53–59}, numpages = {7}, keywords = {benchmark dataset, debugging, language models, socratic dialogue}, location = {<conf-loc>, <city>Portland</city>, <state>OR</state>, <country>USA</country>, </conf-loc>}, series = {SIGCSE 2024} } ```
许可证:MIT协议 任务类别: - 文本到文本生成 - 文本生成 语言:英语 标签: - 代码 数据集名称:苏格拉底式调试基准(Socratic Debugging Benchmark) --- # 苏格拉底式调试基准(Socratic Debugging Benchmark) <p align="center"> <img src="img/socratic_debugging.png" height="400"> </p> 本仓库包含配套两篇学术论文的苏格拉底式调试基准(Socratic Debugging Benchmark)数据集:一篇为发表于ACL 2023会议的第18届“自然语言处理创新应用于教育构建”研讨会(18th Workshop on Innovative Use of NLP for Building Educational Applications, BEA 2023)的《针对新手调试者的苏格拉底式提问:基准数据集与初步评估》(Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations),另一篇为发表于SIGCSE'24的《语言模型能否运用苏格拉底式方法?代码调试相关实验》(Can Language Models Employ the Socratic Method? Experiments with Code Debugging)。 该数据集同时托管于[GitHub](https://github.com/taisazero/socratic-debugging-benchmark)平台。 若使用本数据集,请引用如下论文: @inproceedings{al-hossami-etal-2023-socratic, title = "Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations", author = "Al-Hossami, Erfan and Bunescu, Razvan and Teehan, Ryan and Powell, Laurel and Mahajan, Khyati and Dorodchi, Mohsen", booktitle = "Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023)", month = jul, year = "2023", address = "Toronto, Canada", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.bea-1.57", pages = "709--726", abstract = "苏格拉底式提问是一种教学策略,即引导学生自主解决问题,而非直接给出答案。本文构建了一套苏格拉底式对话数据集,其中指导者协助新手程序员修复简单计算问题的带缺陷代码解决方案。该数据集被用于评估基于GPT的大语言模型(Large Language Model, LLM)的苏格拉底式调试能力。实验发现,GPT-4的表现远优于GPT-3.5,但其精确率与召回率仍不及人类专家,这为该领域的后续研究提供了改进方向。", } @inproceedings{al-hossami-etal-2024-can, author = {Al-Hossami, Erfan and Bunescu, Razvan and Smith, Justin and Teehan, Ryan}, title = {Can Language Models Employ the Socratic Method? Experiments with Code Debugging}, year = {2024}, isbn = {9798400704239}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3626252.3630799}, doi = {10.1145/3626252.3630799}, abstract = "在苏格拉底式教学方法中,教师会引导学生自主解决问题,而非直接提供解决方案。该策略虽能显著提升学习效果,但通常耗时且认知成本高昂。自动化苏格拉底式AI智能体(AI Agent)可辅助人类教学并扩大覆盖规模,但其发展却因缺乏适用于训练与评估的合适数据集而受阻。本文构建了一套人工标注的多轮苏格拉底式对话数据集,用于协助新手程序员修复简单计算问题的带缺陷代码解决方案。该数据集被用于评估多款大语言模型的苏格拉底式调试能力,包括对基于指令的文本到文本Transformer模型Flan-T5进行微调,以及对超大规模的GPT-4进行零样本(Zero-shot)与思维链(Chain-of-Thought)提示。本研究的代码与数据集已在下方链接面向研究人员免费开放。", booktitle = {Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1}, pages = {53–59}, numpages = {7}, keywords = {benchmark dataset, debugging, language models, socratic dialogue}, location = {<conf-loc>, <city>Portland</city>, <state>OR</state>, <country>USA</country>, </conf-loc>}, series = {SIGCSE 2024} }
Socratic Debugging Benchmark 数据集概述
基本信息
- 许可证:MIT
- 任务类别:
- 文本到文本生成
- 文本生成
- 语言:英语
- 标签:代码
- 美观名称:Socratic Debugging Benchmark
数据集描述
该数据集伴随以下论文发布:
- "Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations",收录于 ACL 2023 的第18届NLP在教育应用创新研讨会(BEA 2023)。
- "Can Language Models Employ the Socratic Method? Experiments with Code Debugging",收录于 SIGCSE24 会议。
引用信息
如果您使用此数据集,请引用以下论文:
@inproceedings{al-hossami-etal-2023-socratic, title = "Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations", author = "Al-Hossami, Erfan and Bunescu, Razvan and Teehan, Ryan and Powell, Laurel and Mahajan, Khyati and Dorodchi, Mohsen", booktitle = "Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023)", month = jul, year = "2023", address = "Toronto, Canada", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.bea-1.57", pages = "709--726", abstract = "Socratic questioning is a teaching strategy where the student is guided towards solving a problem on their own, instead of being given the solution directly. In this paper, we introduce a dataset of Socratic conversations where an instructor helps a novice programmer fix buggy solutions to simple computational problems. The dataset is then used for benchmarking the Socratic debugging abilities of GPT-based language models. While GPT-4 is observed to perform much better than GPT-3.5, its precision, and recall still fall short of human expert abilities, motivating further work in this area.", }
@inproceedings{al-hossami-etal-2024-can, author = {Al-Hossami, Erfan and Bunescu, Razvan and Smith, Justin and Teehan, Ryan}, title = {Can Language Models Employ the Socratic Method? Experiments with Code Debugging}, year = {2024}, isbn = {9798400704239}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3626252.3630799}, doi = {10.1145/3626252.3630799}, abstract = {When employing the Socratic method of teaching, instructors guide students toward solving a problem on their own rather than providing the solution directly. While this strategy can substantially improve learning outcomes, it is usually time-consuming and cognitively demanding. Automated Socratic conversational agents can augment human instruction and provide the necessary scale, however their development is hampered by the lack of suitable data for training and evaluation. In this paper, we introduce a manually created dataset of multi-turn Socratic advice that is aimed at helping a novice programmer fix buggy solutions to simple computational problems. The dataset is then used for benchmarking the Socratic debugging abilities of a number of language models, ranging from fine-tuning the instruction-based text-to-text transformer Flan-T5 to zero-shot and chain of thought prompting of the much larger GPT-4. The code and datasets are made freely available for research at the link below.}, booktitle = {Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1}, pages = {53–59}, numpages = {7}, keywords = {benchmark dataset, debugging, language models, socratic dialogue}, location = {<conf-loc>, <city>Portland</city>, <state>OR</state>, <country>USA</country>, </conf-loc>}, series = {SIGCSE 2024} }




