Touch-Vision-Language (TVL) Dataset
收藏资源简介:
TVL数据集由加州大学伯克利分校创建,包含44,000对视觉-触觉配对数据,其中10%由人工标注,90%由GPT-4V生成伪标签。该数据集旨在解决多模态对齐问题,特别是视觉、触觉和语言之间的对齐。数据集通过特制的3D打印设备在野外同步收集触觉和视觉数据,用于训练视觉-语言对齐的触觉编码器和触觉-视觉-语言模型,以提高多模态理解和文本生成的能力。
The TVL dataset, created by the University of California, Berkeley, consists of 44,000 pairs of visual-tactile paired data. Of these, 10% are manually annotated, while the remaining 90% have pseudo-labels generated by GPT-4V. This dataset is designed to address multimodal alignment issues, particularly the alignment between vision, touch and language. The tactile and visual data were synchronously collected in field environments using a purpose-built 3D printing device. It is used to train vision-language aligned tactile encoders and tactile-vision-language models, so as to enhance multimodal understanding and text generation capabilities.

- 1A Touch, Vision, and Language Dataset for Multimodal Alignment加州大学伯克利分校 · 2024年



