ADVLAT-Engine
收藏资源简介:
ADVLAT-Engine是一个用于自动生成大量指令-动作对数据集的自动化数据收集原型系统。该数据集由加州大学默塞德分校的Mi3实验室的研究团队创建,旨在通过使用GPS应用程序和自然语言处理技术,自动收集和分类各种指令,并配合视频数据形成完整的视觉-语言-动作三元组。该数据集包含来自Google Maps、Apple Maps和Waze等导航应用程序的指令,并按照不同的分类进行标注。ADVLAT-Engine可以自动收集数据,包括视频(视觉)、指令(语言)和车辆轨迹(动作),用于训练自主视觉语言导航模型。该数据集的创建过程涉及使用GPS应用程序收集指令,并使用OpenAI Whisper模型进行语音转录,然后将指令与视频帧和GPS位置同步。ADVLAT-Engine的应用领域包括视觉语言导航和人机交互自主系统,旨在解决数据集创建过程中人力成本高、效率低的问题。
ADVLAT-Engine is an automated data collection prototype system for automatically generating large-scale instruction-action pair datasets. This dataset was created by the research team from the Mi3 Lab at the University of California, Merced. It aims to automatically collect and categorize various instructions using GPS applications and natural language processing technologies, and combine them with video data to form complete vision-language-action triplets. This dataset contains instructions from navigation applications such as Google Maps, Apple Maps, and Waze, and is annotated with different categories. ADVLAT-Engine can automatically collect data including video (vision), instructions (language), and vehicle trajectories (action) for training autonomous vision-language navigation models. The creation process of this dataset involves collecting instructions via GPS applications, using the OpenAI Whisper model for speech transcription, and then synchronizing the instructions with video frames and GPS locations. The application scenarios of ADVLAT-Engine include vision-language navigation and human-computer interaction autonomous systems, aiming to address the issues of high labor costs and low efficiency during dataset creation.




