TLV (Touch-Language-Vision)
收藏资源简介:
TLV数据集是由北京交通大学北京交通数据分析与挖掘重点实验室创建,旨在通过人机协同方式实现触觉、语言和视觉的多模态对齐。该数据集包含20000对同步的触觉和视觉观察,其中详细标注了19834个实例,每个实例都附有句子级别的描述。创建过程中,通过使用GelSight传感器和VisGel数据集收集触觉和视觉数据,并利用GPT-4V进行文本标注。TLV数据集主要应用于触觉相关的多模态感知研究,特别是解决触觉与语言之间的语义对齐问题,为机器人和人工智能领域提供了丰富的研究资源。
The TLV Dataset was developed by the Key Laboratory of Beijing Traffic Data Analysis and Mining, Beijing Jiaotong University, aiming to achieve multimodal alignment of tactile, linguistic and visual modalities through human-machine collaboration. This dataset includes 20,000 pairs of synchronized tactile and visual observations, with 19,834 instances meticulously annotated, each attached with a sentence-level description. During the dataset construction, tactile and visual data were collected using the GelSight sensor and the VisGel dataset, and text annotations were generated via GPT-4V. The TLV Dataset is primarily applied in tactile-related multimodal perception research, particularly to solve the semantic alignment problem between touch and language, providing abundant research resources for the robotics and artificial intelligence fields.




