MECAT
收藏资源简介:
MECAT是一个多专家构建的细粒度音频理解任务基准数据集,由MiLM Plus和小米集团的研究人员创建。该数据集包含约20,000个音频剪辑,涵盖了八个不同的音频领域,包括纯音域(如寂静、语音、声音事件和音乐)以及所有可能的混合音域。数据集提供了丰富的标注,包括细粒度的音频描述和开放式的问答对,旨在评估模型在复杂音频场景下的理解能力。MECAT的创建过程结合了专门的专家模型和大型语言模型的推理,以提供多角度、细粒度的描述和开放式的问答对。数据集的应用领域包括音频描述、音频问答等,旨在解决现有基准数据集在评估音频理解方面的局限性,提高模型的感知准确性和描述细节。
MECAT is a fine-grained audio understanding task benchmark dataset constructed by multiple experts, developed by researchers from MiLM Plus and Xiaomi Group. This dataset contains approximately 20,000 audio clips, covering eight distinct audio domains, including single-modality audio categories (such as silence, speech, sound events, and music) as well as all possible mixed audio domains. It provides rich annotations including fine-grained audio descriptions and open-ended question-answer pairs, aiming to evaluate models' understanding capabilities in complex audio scenarios. The construction process of MECAT integrates inference from specialized expert models and Large Language Models (LLMs) to generate multi-perspective, fine-grained descriptions and open-ended question-answer pairs. The applicable scenarios of this dataset include audio captioning, audio question answering and other related tasks, which is designed to address the limitations of existing benchmark datasets in audio understanding evaluation and improve the perceptual accuracy and detail-description performance of models.

- 1MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks小米集团,中国北京; 香港中文大学,中国香港 · 2025年



