VISOR(EPIC-KITCHENSVISOR)
收藏资源简介:
我们很自豪地宣布EPIC-KITCHENS遮阳板,一个新的像素注释数据集,以及一个基准套件,用于分割以自我为中心的视频中的手和活动对象。VISOR注释来自EPIC-KITCHENS的视频,这伴随着当前视频分割数据集中未遇到的一系列新挑战。具体来说,我们需要确保像素级注释的短期和长期一致性,因为对象经历了变革性的相互作用,例如洋葱被去皮,切丁和煮熟-我们的目标是获得果皮,洋葱片,砧板,刀,锅,以及演戏的手。VISOR引入了一个注释管道,部分由人工智能驱动,用于可扩展性和质量,并介绍了:
We are proud to announce EPIC-KITCHENS VISOR, a novel pixel-wise annotated dataset and a benchmark suite for segmenting hands and active objects in egocentric videos. VISOR annotations are derived from videos in the EPIC-KITCHENS corpus, which introduces a set of novel challenges unaddressed in existing video segmentation datasets. Specifically, it is critical to ensure short-term and long-term consistency of pixel-level annotations when objects undergo transformative interactions, such as an onion being peeled, diced, and cooked. Our annotation targets include peels, onion slices, cutting boards, knives, pots, and the hands executing these actions. VISOR presents an annotation pipeline partially driven by AI to guarantee scalability and annotation quality, and introduces:




