Full-Duplex-Bench-v3 (FDB-v3)
收藏资源简介:
Full-Duplex-Bench-v3(FDB-v3)是由台湾大学和英伟达联合创建的语音交互评估数据集,旨在测试语音代理在真实场景中的多步骤工具使用能力。该数据集包含100条真实人类语音记录,覆盖了五种不同的语言不流畅现象(如填充词、停顿、犹豫等),并标注了四种任务领域的多步骤API调用场景。数据来源于12名不同背景的说话者在非受控环境下的自然语音采集,通过系统性的难度分级和确定性API设计确保评估的可靠性。该数据集主要应用于语音代理的实时交互、工具调用和状态回滚等研究领域,旨在解决语音代理在真实世界应用中面临的语言不流畅和多步骤任务执行难题。
Full-Duplex-Bench-v3 (FDB-v3) is a speech interaction evaluation dataset co-developed by National Taiwan University and NVIDIA, aiming to evaluate the multi-step tool use capabilities of speech agents in real-world scenarios. This dataset includes 100 real human speech recordings, covering five distinct types of speech disfluencies such as filled pauses, silent pauses and hesitations, and is annotated with multi-step API call scenarios across four task domains. The data is collected from natural speech of 12 speakers with diverse backgrounds in uncontrolled environments. The reliability of the evaluation is guaranteed through systematic difficulty grading and deterministic API design. This dataset is mainly applied to research fields including real-time interaction, tool invocation and state rollback of speech agents, and targets to address the challenges of speech disfluencies and multi-step task execution faced by speech agents in real-world applications.




