ADL-X
收藏资源简介:
ADL-X是由北卡罗来纳大学夏洛特分校和英智†蔚蓝海岸大学合作创建的多视角RGBD指令ADL数据集,包含100,000个未修剪的RGB视频-指令对,3D姿态,语言描述和动作条件对象轨迹。该数据集通过新颖的半自动化框架生成,旨在训练大型语言视觉模型(LLVMs)以理解和预测日常生活中的活动。ADL-X的创建过程涉及从NTU RGB+D 120数据集中提取和处理视频,采用人物中心裁剪策略和动作序列随机组合,以捕捉ADL场景中的自然随机性。数据集的应用领域包括老年人护理监控、认知衰退评估和机器人辅助开发,旨在通过精确的时空关系理解和复杂的人机交互来解决现实世界中的问题。
ADL-X is a multi-view RGBD instruction ADL dataset co-developed by the University of North Carolina at Charlotte and Université Côte d'Azur. It contains 100,000 untrimmed RGB video-instruction pairs, 3D poses, linguistic descriptions, and action-conditioned object trajectories. Generated via a novel semi-automated framework, this dataset is designed to train Large Language-Vision Models (LLVMs) for understanding and predicting daily-life activities. The development process of ADL-X involves extracting and processing videos from the NTU RGB+D 120 dataset, adopting person-centric cropping strategies and random combinations of action sequences to capture the natural randomness in ADL scenarios. The application areas of this dataset include elder care monitoring, cognitive decline assessment, and robotic assistance development, aiming to solve real-world problems through accurate spatio-temporal relationship understanding and complex human-computer interaction.

- 1LLAVIDAL: Benchmarking Large Language Vision Models for Daily Activities of Living北卡罗来纳大学夏洛特分校†英智†蔚蓝海岸大学 · 2024年



