遇见数据集

多人场景下的多手手掌检测数据

收藏
浙江省数据知识产权登记平台2025-12-23 更新2025-12-24 收录
官方服务:

资源简介:

该数据集包含大量图像和视频帧,其中捕捉了多个个体在同一画面内进行互动或独立活动的场景,并对所有出现的手掌进行了标注。此数据适用于会议分析系统、智能监控、多人协作的VR/AR环境以及零售业的消费者行为分析等。利用该数据训练的模型能够同时检测并区分画面中的多个手掌,解决了因手部交叉、重叠和数量变化导致的检测混淆与失败问题,为理解群体互动行为提供了基础技术支持。在多人场景下准确检测所有手掌是实现群体行为分析的前提。具体过程包括:(1)数据收集:采集包含多个手掌的图像,并为每个手掌实例标注唯一的标注边界框。(2)数据处理:采用数据增强技术(随机缩放、裁剪、旋转)以增加模型对多手场景复杂性的适应能力。多人场景图像特征表示图像在高维特征空间的映射,通过公式 F_multi-hand​=Encodercnn​(I_scene) 提取,其中 F_multi-hand为多人场景图像特征,I_scene为输入的多人场景图像,Encodercnn为预训练的卷积神经网络。(3)模型构建:搭建一个能够输出多个目标边界框的检测模型。根据公式 pred_bbox,conf=Decodermulti-det(F_multi-hand) 从场景特征中解码出所有手掌的边界框集合,其中pred_bbox为预测边界框,conf为预测置信度,Decodermulti-det表示目标检测模型;关键评估指标为平均精度均值(mean Average Precision, mAP),用于综合评估模型多个手掌的综合性能。

This dataset contains a large number of images and video frames, capturing scenarios where multiple individuals interact or perform independent activities within the same frame, with all visible hands annotated. This data is applicable to scenarios such as conference analysis systems, intelligent surveillance, multi-person collaborative VR/AR environments, and consumer behavior analysis in the retail industry. Models trained with this dataset can simultaneously detect and distinguish multiple hands in a frame, solving the detection confusion and failures caused by hand crossing, overlapping, and varying quantities, providing foundational technical support for understanding group interactive behaviors. Accurately detecting all hands in multi-person scenarios is the prerequisite for group behavior analysis. The specific process includes: (1) Data Collection: Collect images containing multiple hands, and annotate each hand instance with a unique bounding box. (2) Data Processing: Adopt data augmentation techniques (random scaling, cropping, and rotation) to enhance the model's adaptability to the complexity of multi-hand scenarios. The image feature of multi-person scenarios represents the mapping of images in the high-dimensional feature space, which is extracted via the formula $F_{ ext{multi-hand}} = ext{Encoder}_{ ext{cnn}}(I_{ ext{scene}})$, where $F_{ ext{multi-hand}}$ is the image feature of multi-person scenarios, $I_{ ext{scene}}$ is the input multi-person scenario image, and $ ext{Encoder}_{ ext{cnn}}$ is a pre-trained Convolutional Neural Network. (3) Model Construction: Build a detection model capable of outputting multiple target bounding boxes. The set of all hand bounding boxes is decoded from the scene features via the formula $ ext{pred\_bbox}, ext{conf} = ext{Decoder}_{ ext{multi-det}}(F_{ ext{multi-hand}})$, where $ ext{pred\_bbox}$ is the predicted bounding box, $ ext{conf}$ is the prediction confidence, and $ ext{Decoder}_{ ext{multi-det}}$ represents the object detection model. The key evaluation metric is mean Average Precision (mAP), which is used to comprehensively evaluate the model's performance on multi-hand detection tasks.

创建时间:
2025-10-14
搜集汇总
数据集介绍
多人场景下的多手手掌检测数据 数据集图片
背景与挑战
背景概述
该数据集专注于多人场景下的多手手掌检测,包含图像和视频帧数据,并对所有手掌进行了标注,适用于会议分析、智能监控和VR/AR等场景。它通过数据增强和特征提取技术,支持模型同时检测多个手掌,解决手部交叉和重叠导致的检测问题,并以平均精度均值(mAP)作为关键评估指标。数据集规模为65.99条,格式为zip,由企业自行产生,按需更新,为群体互动行为分析提供基础技术支持。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务