VLT-MI
收藏资源简介:
VLT-MI数据集由中国科学院自动化研究所创建,是首个支持多轮多模态交互的视觉语言跟踪基准。该数据集包含3619个视频,总帧数达到660万,涵盖短期、长期和全局实例跟踪任务。数据集通过DTLLM-VLT生成高质量的视频文本信息,实现动态交互,旨在解决传统VLT基准在多轮交互中的不足。VLT-MI的应用领域主要集中在视觉语言跟踪任务,旨在通过多模态交互提升跟踪器的准确性和鲁棒性。
The VLT-MI dataset, created by the Institute of Automation, Chinese Academy of Sciences, is the first visual-language tracking benchmark supporting multi-turn multimodal interactions. It consists of 3,619 videos with a total of 6.6 million frames, covering short-term, long-term and global instance tracking tasks. The dataset generates high-quality video-text information through DTLLM-VLT to enable dynamic interactions, aiming to address the limitations of traditional visual-language tracking (VLT) benchmarks in multi-turn interaction scenarios. The application domains of VLT-MI primarily focus on visual-language tracking tasks, with the objective of improving the accuracy and robustness of trackers via multimodal interactions.




