遇见数据集

DriverEmo

收藏
DataCite Commons2025-12-16 更新2026-02-09 收录
官方服务:

资源简介:

Driver emotional states significantly influence road safety, yet comprehensive benchmarks for emotion-aware driving systems remain limited. This paper presents the first large-scale multimodal benchmark for simultaneous driver emotion recognition, behavior classification, and driving scene understanding using Vision-Language Models (VLMs). We introduce a dataset comprising 13,200 visual question-answering pairs and 2,200 dense captions, evaluating state-of-the-art VLMs across three dimensions: driver emotion recognition across six emotional states, behavior classification spanning eight types, and contextual scene understanding including weather, time of day, traffic density, and vehicle motion. Our evaluation of multiple VLMs reveals distinct performance patterns: while models demonstrate strong capabilities in identifying environmental context such as weather and time of day, they exhibit significant limitations in recognizing driver emotions and describing vehicle dynamics. These findings highlight fundamental challenges in applying current VLMs to human-centric and dynamics-aware driving scenarios. This benchmark establishes a standardized evaluation framework for emotion-aware advanced driver assistance systems and provides insights into multimodal reasoning capabilities required for comprehensive driver state monitoring.

驾驶员情绪状态对道路交通安全具有显著影响,但面向情绪感知驾驶系统的全面基准测试仍较为匮乏。本文首次提出基于视觉语言模型(Vision-Language Models)的大规模多模态基准测试集,用于同步开展驾驶员情绪识别、行为分类与驾驶场景理解任务。本次研究构建的数据集包含13200组视觉问答对与2200条稠密字幕,从三个维度对最先进的视觉语言模型展开评估:一是覆盖六种情绪状态的驾驶员情绪识别任务,二是涵盖八种类型的行为分类任务,三是包含天气、时段、交通密度与车辆运动的上下文场景理解任务。对多款最先进视觉语言模型的评估结果呈现出鲜明的性能特征:尽管模型在识别天气、时段等环境上下文方面表现出色,但在识别驾驶员情绪与描述车辆动态层面存在显著局限。上述研究结果凸显了将当前视觉语言模型应用于以人为中心、且需感知动态信息的驾驶场景时所面临的根本性挑战。该基准测试集为面向情绪感知的高级驾驶辅助系统搭建了标准化评估框架,同时为全面驾驶员状态监测所需的多模态推理能力提供了重要研究启示。

提供机构:
figshare
创建时间:
2025-12-16
搜集汇总
数据集介绍
DriverEmo 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务