Synthetic Clarification and Correction Dialogues
收藏资源简介:
本研究构建了一个名为'Synthetic Clarification and Correction Dialogues'的数据集,该数据集通过教师-学生框架生成,包含针对数据为中心任务的合成澄清和纠正对话。数据集由微软研究院的Christian Poelitz和Nick McKenna开发,旨在解决表格数据问题解答中的信息不完整问题。数据集通过模拟AI助手与用户之间的多轮对话,包含AI发起的澄清和用户发起的纠正两种场景。这些对话是从现有的数据集中生成的,包含完整的表格问答示例。数据集的创建过程涉及对现有数据集的信息消减,然后通过教师模型指导学生模型生成澄清问题和进行纠正。该数据集的应用领域是数据为中心的任务,特别是在表格数据问题解答中,旨在提高AI模型在面对不完整信息时的处理能力。
This study constructs a dataset named 'Synthetic Clarification and Correction Dialogues'. Generated via a teacher-student framework, this dataset contains synthetic clarification and correction dialogues tailored for data-centric tasks. Developed by Christian Poelitz and Nick McKenna from Microsoft Research, this dataset aims to address the problem of incomplete information in table data question answering. The dataset simulates multi-turn dialogues between AI assistants and users, encompassing two scenarios: AI-initiated clarification and user-initiated correction. These dialogues are generated from existing datasets and include complete table question-answering examples. The dataset creation process involves information reduction on existing datasets, followed by the generation of clarification questions and corrections by student models guided by teacher models. The application scope of this dataset covers data-centric tasks, particularly in table data question answering, with the goal of enhancing the capability of AI models to handle incomplete information.

- 1Synthetic Clarification and Correction Dialogues about Data-Centric Tasks -- A Teacher-Student Approach微软研究院 · 2025年



