AVQA-Hard, Music-AVQA-Hard
收藏资源简介:
AVQA-Hard 和 Music-AVQA-Hard 是针对现有视频问答基准测试中存在的视觉捷径问题进行筛选后的数据集。这些数据集旨在强调音频信息的重要性,并推动视频问答模型在音频理解和时序整合方面的研究。数据集的构建过程涉及对现有视频问答数据集进行单帧推理测试,并筛选出那些不能仅通过视觉信息就能解决的项目,从而构成对音频信息敏感的难题集。这些数据集的应用领域在于评估和改进视频问答模型的音频理解和时序整合能力,以期更好地反映实际应用场景中用户对视频内容的理解需求。
AVQA-Hard and Music-AVQA-Hard are datasets curated to mitigate the visual shortcut problems prevalent in existing video question answering (VideoQA) benchmarks. These datasets are designed to emphasize the critical role of audio information and promote research on audio comprehension and temporal integration for video QA models. The construction process entails performing single-frame inference tests on existing video QA datasets, and screening out instances that cannot be resolved using only visual cues, thereby forming a challenging problem set sensitive to audio information. These datasets serve to evaluate and enhance the audio understanding and temporal integration capabilities of video QA models, thereby better aligning with users' demands for comprehending video content in real-world application scenarios.
LLaVA-AV-SSM 数据集概述
基本信息
- 数据集名称: LLaVA-AV-SSM
- 研究主题: 音频对现代视频大语言模型及其基准测试的重要性
- 状态: 预印本(arXiv:2509.17901),正在评审中
- 发布日期: 2025年9月22日
核心发现
- 现有视频理解基准测试大多可通过单帧图像解决,音频常被忽略且不影响性能
- 在标准测试集上音频带来的提升有限
- 在音频敏感的子集(AVQA-Hard、Music-AVQA-Hard)上音频起决定性作用
技术方案
- 基于LLaVA架构,增加语音/音频编码器
- 使用轻量级Mamba状态空间压缩器解决音频令牌爆炸问题
- 提出音频敏感评估方法
评估基准
- 主流测试套件: 音频贡献较小
- 定制化子集:
- AVQA-Hard
- Music-AVQA-Hard
相关资源
- 论文地址: https://arxiv.org/abs/2509.17901
- 作者: Geewook Kim, Minjoon Seo

- 1Does Audio Matter for Modern Video-LLMs and Their Benchmarks?NAVER Cloud AI, KAIST AI · 2025年



