TNLLT
收藏资源简介:
TNLLT是一个大规模的长时视觉-语言跟踪基准数据集,包含200个视频序列,其中150个用于训练,50个用于测试。数据集由研究人员从视频网站收集,主要涉及电视、电影、游戏、娱乐、纪录片等。研究人员对裁剪的视频进行了细致的矩形框标注,并标注了对象的视觉外观、运动和其他线索的语言描述。这些视频涵盖了与视觉-语言跟踪相关的15个挑战因素。该数据集为视觉-语言跟踪任务的研究提供了一个坚实的基础。
TNLLT is a large-scale long-term visual-language tracking benchmark dataset consisting of 200 video sequences, with 150 allocated for training and 50 for testing. The dataset was collected by researchers from video websites, mainly covering content from TV series, movies, games, entertainment, documentaries and other categories. Researchers conducted detailed bounding box annotations on the cropped video clips, alongside linguistic descriptions of the target's visual appearance, motion and other relevant cues. These videos encompass 15 challenging factors related to visual-language tracking tasks. This dataset provides a solid foundation for research on visual-language tracking tasks.




