遇见数据集

West Point Heroico Spanish Speech

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<h3>Introduction</h3> <p>This file contains documentation on West Point Heroico Spanish Speech, Linguistic Data Consortium (LDC) catalog number LDC2006S37 and ISBN 1-58563-391-7. </p><p> West Point Heroico Spanish Speech is a database of digital recordings of spoken Spanish. It was designed and collected by staff and faculty of the Department of Foreign Languages (DFL) and Center for Technology Enhanced Language Learning (CTELL) to develop acoustic models for speech recognition systems. The U.S. government uses these systems to provide speech-recognition enhanced language learning courseware to government linguists and students enrolled in various government language programs. Additionally, parts of this corpus were designed to model question/answer dialogues for use in domain-specific speech-to-speech translation systems. The corpus consists of two subcorpora, one collected in September 2001 at El Heroico Colegio Militar (HEROICO), the Mexican Military Academy in Mexico City, and the other at USMA at different times since 1997. The USMA subcorpus includes data from non-native speakers and data collected through a throat microphone.</p> <h3>Data</h3> <p> Two kinds of prompt scripts were used, one to elicit read speech and one for free-response answers to questions. The read speech prompts are also divided into two groups, one designed to elicit speech typical of language learning scenarios and the other for speech from educated native speakers. The scripts used to record read speech have a total of 724 distinct sentences. This number includes 205 short, simple sentences used in typical language learning scenarios. The other 519 sentences were extracted from lecture notes used at USMA in a military readings course. All of the read speech prompts are listed in two files in the transcripts directory: HEROICO- Recordings.txt and USMA-prompts.txt, containing the sentences read by informants at the Mexican Military Academy and USMA, respectively. Each line of these files has two fields separated by a tab, the first denoting the base name of the waveform file, and the second the prompt used in recording the utterence.</p> <p> The read speech data collected from informants at HEROICO are stored in the HEROICO/Recordings Spanish directory. The script used to elicit free-response answers contains 143 questions. The text that was actually presented to the informants is in the file named questions.txt in the transcripts directory. Data recorded from these prompts are stored in the HEROICO/Answers Spanish directory. The human-performed transcriptions of the informants answers are listed in the HEROICO-Answers.txt file in the transcripts directory. Again, each line of this file has two fields separated by a tab the first field contains two numbers separated by a slash. The first number is an identification index for the speaker. The second number is an index to the question. The second field on the line contains a word level transcription of the informants answer to the question indexed by the second number in the first field. So for example in the line: 100/10 no ella no tiene barba ni bigote no ella no tiene barba ni bigote is a transcription of the response speaker 100 gave to question 10. The corresponding waveform file is stored in the file 10.wav in the directory HEROICOAnswers Spanish100. Each speaker in the HEROICO subcorpus attempted to record 100 utter- ances by reading 75 sentences and giving 25 free-response answers to questions.</p> <p> Both native and non-native USMA informatnts read from the list of 205 simple sentences. The prompts used in the USMA subcorpus are listed in the file USMA-prompts.txt in the transcripts directory. This file has the same two-field format as the above transcription files. Some of the USMA informants wore an additional throat microphone. That data was recorded in a separate stream and stored in files whose names begin with the letter t. Data collected at USMA are stored under the USMA directory. The names of the directories under the USMA directory indicate whether the speaker was native or non-native. The speakers native country is also indicated in the case of native speakers.</p> <p> Speech data was collected at HEROICO using Pentium 450 mHz laptop computers running Windows 2000 with a 16-bit data size and sampling rate of 22,050 Hz. The recording script presented a visual display of the sentence to be recorded. The informant pressed a key and spoke the sentence. The recording was played back for review allowing the utterance to be re- recorded. A member of the data collection team was on hand during the recording session to verify recordings and provide technical assistance in case of malfunctioning equipment.</p> <p> The data from USMA was collected using several different microphones and formats. Most of the data were recorded on Pentium computers running Linux through an m-10 Shuer head-mounted microphone. Entropics ESPS programs were used in most cases, especially when both head-mounted and throat microphones were used.</p> <h3>Samples</h3> <p>For an example of the data in this corpus, please listen to this <a href="./desc/addenda/LDC2006S37.wav" rel="nofollow">audio sample</a>. </p> </br> Portions © 2001 United States Military Academy, © 2006 Trustees of the University of Pennsylvania

<h3>简介</h3> <p>本文件包含关于西点英雄学院西班牙语语音语料库(West Point Heroico Spanish Speech)的说明文档,其语言数据联盟(Linguistic Data Consortium, LDC)目录号为LDC2006S37,ISBN为1-58563-391-7。</p><p> 西点英雄学院西班牙语语音语料库是一个西班牙语口语数字录音数据库,由外语系(Department of Foreign Languages, DFL)与技术增强语言学习中心(Center for Technology Enhanced Language Learning, CTELL)的教职员工设计并采集,旨在为语音识别系统开发声学模型。美国政府利用此类系统为政府语言学家以及参与各类政府语言项目的学员提供搭载语音识别功能的语言学习课件。此外,该语料库的部分内容旨在构建问答对话模型,用于特定领域的语音到语音翻译系统。该语料库包含两个子语料库:其一于2001年9月在墨西哥城的墨西哥军事学院埃尔英雄科莱吉奥米利塔尔(El Heroico Colegio Militar, HEROICO)采集;其二自1997年起于美国军事学院(United States Military Academy, USMA)分多次采集。USMA子语料库包含非母语使用者的语音数据,以及通过喉头麦克风采集的数据。</p> <h3>数据</h3> <p> 本次采集使用了两类提示脚本:一类用于引导朗读语音,另一类用于生成针对问题的自由应答语音。朗读语音提示又分为两组:一组用于采集典型语言学习场景下的语音,另一组用于采集受过良好教育的母语使用者的语音。用于录制朗读语音的脚本共包含724个独特句子,其中包括205个用于典型语言学习场景的简短简单句,剩余519个句子取自美国军事学院(USMA)军事阅读课程的讲义。所有朗读语音提示均收录于转录目录下的两个文件中:HEROICO-Recordings.txt与USMA-prompts.txt,分别包含墨西哥军事学院与USMA的受访对象朗读的句子。上述文件的每一行均包含两个以制表符分隔的字段:第一个字段为波形文件的基名,第二个字段为录制该话语时使用的提示文本。</p> <p> 从HEROICO受访对象处采集的朗读语音数据存储于HEROICO/Recordings Spanish目录中。用于引导自由应答的脚本包含143个问题,实际向受访对象展示的文本收录于转录目录下的questions.txt文件中。基于此类提示录制的数据存储于HEROICO/Answers Spanish目录中。受访对象应答的人工转录文本收录于转录目录下的HEROICO-Answers.txt文件中。该文件的每一行同样包含两个以制表符分隔的字段:第一个字段包含两个以斜杠分隔的数字,前者为说话者的标识索引,后者为问题的索引;第二个字段为该说话者针对对应问题的应答的词级转录文本。例如某行内容为:100/10 no ella no tiene barba ni bigote,其中no ella no tiene barba ni bigote即为说话者100针对问题10的应答转录文本。对应的波形文件存储于HEROICO/Answers Spanish/100目录下的10.wav文件中。HEROICO子语料库中的每位受访对象需录制100条话语:朗读75个句子,并针对25个问题给出自由应答。</p> <p> USMA的母语与非母语受访对象均需朗读205个简单句列表。USMA子语料库中使用的提示文本收录于转录目录下的USMA-prompts.txt文件中,该文件采用与上述转录文件相同的双字段格式。部分USMA受访对象额外佩戴了喉头麦克风,此类数据以独立流形式录制,文件名以字母t开头。USMA采集的数据存储于USMA目录下,该目录下的子目录名称可区分说话者为母语使用者还是非母语使用者;对于母语使用者,还会标注其母语国家。</p> <p> HEROICO的语音数据采集使用搭载Windows 2000系统的奔腾450MHz笔记本电脑,采样位数为16位,采样率为22050Hz。录制脚本会可视化展示待录制的句子,受访对象按下按键后朗读该句子,录制完成后会播放录音以供审核,允许重新录制话语。数据采集团队成员会在录制过程中在场,以验证录制效果,并在设备故障时提供技术支持。</p> <p> USMA的数据采集使用了多种不同的麦克风与录制格式。大部分数据通过搭载Linux系统的奔腾电脑,搭配m-10舒尔(Shure)头戴式麦克风录制。多数情况下,尤其是同时使用头戴式麦克风与喉头麦克风时,会使用Entropics ESPS软件进行录制。</p> <h3>示例</h3> <p>如需查看本语料库数据示例,请收听此<a href="./desc/addenda/LDC2006S37.wav" rel="nofollow">音频样本</a>。 </br> 部分内容 © 2001 美国军事学院,© 2006 宾夕法尼亚大学校董会

创建时间:
2020-11-30
搜集汇总
数据集介绍
West Point Heroico Spanish Speech 数据集图片
背景与挑战
背景概述
West Point Heroico Spanish Speech是一个西班牙语语音数据集,包含约19,000个音频文件及对应转录,主要用于开发语音识别系统和语言学习课程。该数据集由墨西哥军事学院和美国军事学院收集,包括阅读和自由回答两种语音类型,并涉及母语和非母语者,采样率为22,050 Hz,适用于语音技术研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务