SQuTR
收藏资源简介:
SQuTR是一个用于评估语音查询到文本检索系统鲁棒性的大规模基准数据集,由多个研究机构联合创建。该数据集整合了来自六个广泛使用的英文和中文文本检索数据集的37,317个独特查询,覆盖金融、多跳问答、开放域问答、医学等多个领域。数据集通过200名真实说话者的语音配置文件合成语音,并在受控的信噪比水平下混合了17类真实环境噪声,从而实现了从安静到高噪声条件下的可重复鲁棒性评估。SQuTR旨在解决现有评估数据集在复杂声学扰动下评估语音查询检索系统鲁棒性不足的问题,为相关研究提供了标准化的测试平台。
Jointly developed by multiple research institutions, SQuTR is a large-scale benchmark dataset for evaluating the robustness of speech query-to-text retrieval systems. It integrates 37,317 unique queries sourced from six widely used English and Chinese text retrieval datasets, covering multiple domains including finance, multi-hop question answering (QA), open-domain QA, and medical research. The dataset synthesizes speech using voice profiles from 200 real speakers, and mixes 17 types of real-world environmental noises at controlled signal-to-noise ratio (SNR) levels, enabling reproducible robustness evaluations across conditions ranging from quiet to high-noise environments. SQuTR aims to resolve the inadequacy of existing evaluation datasets in assessing the robustness of speech query retrieval systems under complex acoustic perturbations, providing a standardized testbed for relevant research in the field.



