COMBO
收藏资源简介:
COMBO数据集由上海科技大学创建,是一个全面的开放知识图谱规范化基准。该数据集包含18000个三元组,不仅提供了实体级名词短语的金标准规范化,还额外提供了关系短语和本体级名词短语的金标准规范化。数据集的构建基于大规模本体知识图谱Wikidata,通过远程监督和人工修订确保数据质量。COMBO数据集的应用领域广泛,旨在解决开放知识图谱中的冗余和歧义问题,提高查询效率和准确性。
The COMBO dataset, created by ShanghaiTech University, is a comprehensive benchmark for open knowledge graph normalization. It contains 18,000 triples, and provides gold-standard normalization not only for entity-level noun phrases, but also for relation phrases and ontology-level noun phrases. Built upon Wikidata, a large-scale ontology knowledge graph, the dataset ensures data quality through distant supervision and manual revision. The COMBO dataset has a wide range of application scenarios, and is designed to address the redundancy and ambiguity issues in open knowledge graphs, thereby improving query efficiency and accuracy.




