Lexicographic Data Retrieval on Knowledge Graphs with SPARQL
收藏资源简介:
该数据集由苏黎世大学信息学院的研究人员创建,旨在为自然语言查询知识图谱中的词典数据提供接口。数据集包含超过120万个映射,将自然语言表达映射到SPARQL查询,为词典数据检索提供模板。该数据集使用GPT2、Phi-1.5和GPT-3.5-Turbo等大型语言模型进行实验,以评估不同模型的能力。数据集的创建过程基于四维分类法,捕捉了Wikidata词典数据本体模块的复杂性。数据集旨在解决知识图谱中词典数据检索的挑战,特别是对于非技术用户来说,它提供了一个更易于使用的接口,以获取结构化的语言知识。
This dataset was developed by researchers from the School of Information, University of Zurich, with the goal of providing an interface for querying dictionary data in knowledge graphs using natural language. It contains over 1.2 million mappings that translate natural language expressions into SPARQL queries, serving as templates for dictionary data retrieval. Experiments have been conducted on this dataset using large language models such as GPT2, Phi-1.5, and GPT-3.5-Turbo to evaluate the performance of different models. The construction of this dataset adopts a four-dimensional classification framework, which captures the complexity of the ontology modules of Wikidata dictionary data. This dataset aims to tackle the challenges of dictionary data retrieval in knowledge graphs; specifically, it offers a more accessible interface for non-technical users to obtain structured linguistic knowledge.




