遇见数据集

<b>D</b>ataset for <b>O</b>pen <b>V</b>ocabulary <b>E</b>ntity <b>G</b>rounding (DOVE-G)

收藏
DataCite Commons2025-06-01 更新2024-08-18 收录
官方服务:

资源简介:

<b>DOVE-G</b>To accommodate the richness of open-vocabulary queries, we introduced a custom dataset—DOVE-G (Dataset for Open-Vocabulary Entity Grounding). This dataset has 8 scenes namely kitchen, kitchenette, room1, room2, room3, bathroom, computer lab, and hallway. This dataset is created to facilitate users to query for objects within a scene using natural language. For each scene within DOVE-G, we manually labeled the ground truth and created 50 natural language queries (<i>L</i><sub><em>q</em></sub> ). To augment this query set, we harnessed LLMs to generate four additional sets of natural language queries. This approach yielded a total of 250 queries for each scene, and cumulatively, we have 4000 queries to evaluate OVSG’s performance. With this setup, we set out to assess how our OVSG framework performs with open-vocabulary queries, one of our key research questions, providing a critical testbed for its effectiveness in handling diverse natural language expressions.

<b>DOVE-G</b>为适配开放词汇查询的丰富性需求,我们构建了定制化数据集DOVE-G(开放词汇实体接地数据集,Dataset for Open-Vocabulary Entity Grounding)。该数据集包含8个场景,分别为厨房、小厨房、房间1、房间2、房间3、浴室、计算机实验室及走廊。构建此数据集的核心目标是支持用户通过自然语言查询场景内的各类物体。针对DOVE-G中的每个场景,我们均人工标注了基准真值(ground truth),并生成了50条自然语言查询语句($L_q$)。为扩充该查询集规模,我们借助大语言模型(Large Language Model,LLM)生成了4组额外的自然语言查询语句。通过该方式,每个场景可获得总计250条查询语句,所有场景累计可提供4000条查询语句,用于评估OVSG框架的性能。基于该数据集构建方案,我们旨在评估OVSG框架在开放词汇查询场景下的表现——这也是我们的核心研究问题之一,同时为该框架在处理多样化自然语言表达时的有效性提供了关键测试平台。

提供机构:
figshare
创建时间:
2023-10-13
搜集汇总
数据集介绍
<b>D</b>ataset for <b>O</b>pen <b>V</b>ocabulary <b>E</b>ntity <b>G</b>rounding (DOVE-G) 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务