EXCGEC
收藏资源简介:
EXCGEC数据集由清华大学等机构创建,专门用于中文语法错误修正任务,包含8216个样本,每个样本都附带有编辑级别的解释。数据集通过半自动化的方法构建,利用GPT-4合成解释并由专业标注人员筛选和优化,确保数据质量。该数据集主要用于提高语法错误修正的解释能力,特别是在教育场景中,帮助学习者理解错误并学习正确的语法规则。
The EXCGEC dataset, created by Tsinghua University and other institutions, is specifically designed for Chinese grammatical error correction (GEC) tasks. It contains 8,216 samples, each accompanied by edit-level explanations. The dataset is constructed via a semi-automated workflow: GPT-4 is used to synthesize the explanations, which are then filtered and optimized by professional annotators to ensure high data quality. This dataset is primarily intended to enhance the explanatory capability of grammatical error correction systems, especially in educational scenarios, to help learners understand errors and master correct grammatical rules.

- 1EXCGEC: A Benchmark of Edit-wise Explainable Chinese Grammatical Error Correction清华大学 · 2024年



