CLC-QuAD
收藏资源简介:
CLC-QuAD是首个基于Wikidata的大规模中文复杂语义解析数据集,由浙江大学等机构创建。该数据集包含超过28,000个问题及其对应的SPARQL查询,涵盖多种问题类型,如事实问题、双重意图问题、布尔问题和计数问题。数据集的构建过程涉及将英文问题翻译成中文,并通过人工验证确保翻译的准确性。CLC-QuAD旨在推动中文知识库问答系统的研究,解决现有数据集在问题类型和语言多样性方面的不足。
CLC-QuAD is the first large-scale Chinese complex semantic parsing dataset based on Wikidata, created by institutions including Zhejiang University. This dataset contains over 28,000 questions and their corresponding SPARQL queries, covering multiple question types such as factoid questions, dual-intent questions, boolean questions, and counting questions. The dataset construction process involves translating English questions into Chinese, with manual verification conducted to ensure translation accuracy. CLC-QuAD aims to promote research on Chinese knowledge base question answering systems, and address the shortcomings of existing datasets in terms of question types and linguistic diversity.

- 1A Chinese Multi-type Complex Questions Answering Dataset over Wikidata浙江大学 · 2021年



