DBpedia NIF
收藏资源简介:
DBpedia NIF是一个大规模、开放且多语言的知识抽取语料库,由德国莱比锡大学的敏捷知识工程与语义网(AKSW) InfAI研究团队创建。该数据集包含了128种语言的维基百科文章内容,旨在深化和扩展DBpedia中的结构化信息,并为各种自然语言处理(NLP)和信息检索(IR)任务提供大规模多语言语言资源。数据集创建过程中,采用了NLP交换格式(NIF)来模型化内容、链接和维基百科文章的信息结构。此外,数据集还通过增加约25%的链接和选择性分区作为链接数据发布而得到进一步丰富。DBpedia NIF的应用领域广泛,包括事实抽取、验证、多语言NLP任务训练等,旨在解决从非结构化文本中提取知识的问题。
DBpedia NIF is a large-scale, open and multilingual knowledge extraction corpus, created by the Agile Knowledge Engineering and Semantic Web (AKSW) InfAI research group at Leipzig University, Germany. This dataset includes Wikipedia article content in 128 languages, with the goal of deepening and expanding the structured information within DBpedia, and providing large-scale multilingual language resources for various natural language processing (NLP) and information retrieval (IR) tasks. During the dataset creation process, the Natural Language Processing Exchange Format (NIF) was adopted to model the information structure of content, links and Wikipedia articles. Furthermore, the dataset has been further enriched by adding approximately 25% additional links and selective partitioning for linked data publication. DBpedia NIF covers a wide range of application fields, including fact extraction, verification, multilingual NLP task training and others, aiming to address the challenge of extracting knowledge from unstructured text.




