遇见数据集

知识图谱查询测试数据集

收藏
官方服务:

资源简介:

面向制造业的知识图谱蕴含了产品设计、制造、装配和服务等流程中的相关知识,具有种类多、数量大、专业强等特点,导致制造业知识图谱中的多模态知识难以被充分利用。同时,不同用户、岗位与任务所需知识具有特异性和动态性。因此,用户难以直接从知识图谱中获取符合自己需求的知识。如何基于多模态制造业知识图谱实现智能化的知识服务,是一个关键问题。为测试指标“意图识别准确率达到90%,智能问答答案的预测精度不低于85%,智能检索目标实体和文档的top5召回率不低于90%。同时,在TeslaV100 GPU上面向10万实体规模的知识图谱平均每次检索时间低于1秒”,从构建的知识图谱中随机选择了部分知识三元组构成数据集。数据内容包括知识图谱中的知识三元组文件,ownthink知识图谱以及相关专利及论文,数据量为8.23GB。

Manufacturing-oriented knowledge graphs contain relevant knowledge across product design, manufacturing, assembly, and service workflows, and are characterized by diverse categories, massive scale, and high professionality, which renders the multimodal knowledge within manufacturing knowledge graphs difficult to be fully exploited. Meanwhile, the knowledge demanded by distinct users, job roles, and tasks exhibits specificity and dynamics. Consequently, users struggle to directly retrieve knowledge tailored to their specific requirements from knowledge graphs. Achieving intelligent knowledge services based on multimodal manufacturing knowledge graphs thus represents a critical research challenge. To validate the following performance metrics: 90% intent recognition accuracy, a prediction accuracy of no less than 85% for intelligent question answering outputs, a top-5 recall rate of no less than 90% for intelligent retrieval of target entities and documents, and an average per-query retrieval time of less than 1 second on a knowledge graph with 100,000 entities when deployed on a TeslaV100 GPU, a subset of knowledge triples was randomly sampled from the constructed knowledge graph to form this dataset. The dataset comprises knowledge triple files extracted from the knowledge graph, the ownthink knowledge graph, as well as relevant patents and academic papers, with a total data size of 8.23 GB.

提供机构:
西安交通大学
搜集汇总
数据集介绍
知识图谱查询测试数据集 数据集图片
背景与挑战
背景概述
该数据集是面向制造业知识图谱的查询测试资源,包含知识三元组、专利及论文等内容,数据量为8.23GB,旨在评估意图识别、智能问答等性能指标。它基于国家重点研发计划项目构建,支持对多模态知识的高效检索与智能化服务测试。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务