遇见数据集

WikiSQE

收藏
arXiv2023-12-30 更新2024-06-21 收录
数据链接:
官方服务:

资源简介:

WikiSQE是一个大规模的数据集,专门用于评估维基百科中句子的质量。该数据集由理化学研究所人工智能和大数据创新研究中心创建,包含了从英文维基百科全历史修订中提取的约340万条句子,每条句子都附有153种质量标签。数据集的创建过程涉及从维基百科的清理模板中手工挑选目标标签,并过滤掉噪声句子。WikiSQE的应用领域广泛,主要用于自然语言处理中的句子质量估计,旨在通过机器学习模型自动检测和分类句子中的质量问题,如引用缺失、语法或语义错误等。

WikiSQE is a large-scale dataset specifically developed for evaluating sentence quality within Wikipedia. Created by the RIKEN Center for Artificial Intelligence and Big Data Innovation, it comprises approximately 3.4 million sentences extracted from the full historical revisions of the English Wikipedia, with each sentence annotated with 153 quality tags. The dataset construction process involves manually selecting target tags from Wikipedia's cleanup templates and filtering out noisy sentences. WikiSQE finds wide applications, primarily serving sentence quality estimation tasks in natural language processing, with the goal of automatically detecting and categorizing quality issues in sentences such as missing citations, grammatical or semantic errors via machine learning models.

创建时间:
2023-05-10
二维码
社区交流群
二维码
科研交流群
商业服务