BAREC-Shared-Task-2026-doc
收藏资源简介:
BAREC-ST-2026数据集是一个用于阿拉伯语细粒度可读性评估的大规模语料库,专为BAREC共享任务2026设计。该数据集包含超过100万词,在句子级别标注了19个可读性等级,并提供了映射到更粗粒度7级、5级和3级分类的方案。文档级别的可读性分数基于其最困难句子的19级可读性等级确定。数据集支持多类别可读性分类任务,涵盖19级、7级、5级和3级分类。数据实例包含ID、文档文件名、句子文本、句子数量、词数、各级可读性标签(如Readability_Level_19、Readability_Level_7等)、来源、书籍、作者、领域(如艺术与人文、STEM、社会科学)和文本类别(如基础、高级、专业)。数据集语言为现代标准阿拉伯语,按文档级别划分为训练集(80%)、开发集(10%)和测试集(10%),并在可读性等级、领域和文本类别上保持平衡。评估任务定义为序数分类,使用准确率(包括19级、7级、5级、3级准确率)、相邻准确率(±1准确率)、平均距离(平均绝对误差)和二次加权Kappa等指标。
The BAREC-ST-2026 dataset is a large-scale corpus for fine-grained readability assessment in Arabic, specifically designed for the BAREC shared task 2026. It contains over 1 million words, annotated at the sentence level with 19 readability grades, and provides mappings to coarser-grained 7-level, 5-level, and 3-level classifications. Document-level readability scores are determined based on the 19-level readability grade of their most difficult sentence. The dataset supports multi-category readability classification tasks, covering 19-level, 7-level, 5-level, and 3-level classifications. Data instances include ID, document filename, sentence text, sentence count, word count, readability labels at various levels (e.g., Readability_Level_19, Readability_Level_7, etc.), source, book, author, domain (such as Arts & Humanities, STEM, Social Sciences), and text category (e.g., basic, advanced, professional). The dataset language is Modern Standard Arabic, divided at the document level into training set (80%), development set (10%), and test set (10%), with balanced distributions across readability grades, domains, and text categories. The evaluation task is defined as ordinal classification, using metrics such as accuracy (including 19-level, 7-level, 5-level, 3-level accuracy), adjacent accuracy (±1 accuracy), average distance (mean absolute error), and quadratic weighted Kappa.




