inesc-id/FalAR
收藏资源简介:
FalAR是一个大规模、带说话人标注的欧洲葡萄牙语语音语料库,基于葡萄牙议会的议会会议录音构建。数据集包含对齐的语音片段、参考转录文本、自动转录文本以及说话人元数据。该发布旨在支持自动语音识别(ASR)、说话人感知语音处理以及欧洲葡萄牙语议会语音相关研究。亮点包括:欧洲葡萄牙语议会语音、约4.9千小时带说话人信息的音频、1,180名说话人附带元数据、覆盖约20年的议会会议、包含参考转录和自动转录、每句话提供字符错误率(CER)以评估转录质量。
FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the Portuguese Parliament. The dataset contains aligned speech segments, reference transcripts, automatic transcripts, and speaker metadata. This release is intended to support research in automatic speech recognition (ASR), speaker-aware speech processing, and related studies on parliamentary speech in European Portuguese. Highlights include: European Portuguese parliamentary speech, ~4.9k hours with speaker information, 1,180 speakers with associated metadata, covers roughly 20 years of parliamentary sessions, includes both a reference transcript and an automatic transcription, and includes a per-utterance character error rate (CER).




