GALE Phase 1 Arabic Broadcast News Parallel Text - Part 2
收藏资源简介:
<h3>Introduction</h3><br> <p>GALE Phase 1 Arabic Broadcast News Parallel Text - Part 2 is the second of the three-part GALE Phase 1 Arabic Broadcast News Parallel Text, which, along with other corpora, was used as training data in year 1 (Phase 1) of the DARPA-funded GALE program. <a href="http://catalog.ldc.upenn.edu/LDC2007T24" rel="nofollow">GALE Phase 1 Arabic Broadcast News Parallel Text - Part 1</a> was released in 2007.</p><br> <p>GALE Phase 1 Arabic Broadcast News Parallel Text - Part 2 contains transcripts and English translations of 10.7 hours of Arabic broadcast news programming selected from various sources. This corpus does not contain the audio files from which the transcripts and translations were generated.</p><br> <p>LDC has released the following GALE Phase 1 & 2 Arabic Parallel Text data sets:</p><br> <ul><br> <li>GALE Phase 1 Arabic Broadcast News Parallel Text - Part 1 (<a href="../../../LDC2007T24">LDC2007T24</a>)</li><br> <li>GALE Phase 1 Arabic Broadcast News Parallel Text - Part 2 (<a href="../../../LDC2008T09">LDC2008T09</a>)</li><br> <li>GALE Phase 1 Arabic Blog Parallel Text (<a href="../../../LDC2008T02">LDC2008T02</a>)</li><br> <li>GALE Phase 1 Arabic Newsgroup Parallel Text - Part 1 (<a href="../../../LDC2009T03">LDC2009T03</a>)</li><br> <li>GALE Phase 1 Arabic Newsgroup Parallel Text - Part 2 (<a href="../../../LDC2009T09">LDC2009T09</a>)</li><br> <li>GALE Phase 2 Arabic Broadcast Conversation Parallel Text Part 1 (<a href="../../../LDC2012T06">LDC2012T06</a>)</li><br> <li>GALE Phase 2 Arabic Broadcast Conversation Parallel Text Part 2 (<a href="../../../LDC2012T14">LDC2012T14</a>)</li><br> <li>GALE Phase 2 Arabic Newswire Parallel Text (<a href="../../../LDC2012T17">LDC2012T17</a>)</li><br> <li>GALE Phase 2 Arabic Broadcast News Parallel Text (<a href="../../../LDC2012T18">LDC2012T18</a>)</li><br> <li>GALE Phase 2 Arabic Web Parallel Text (<a href="../../../LDC2013T01">LDC2013T01</a>)</li><br> </ul><br> <h3>Source Data</h3><br> <p>A total of 10.7 hours of Arabic broadcast news recordings were selected from four sources and four different programs.</p><br> <p>A manual selection procedure was used to choose data appropriate for the GALE program, namely, news and conversation programs focusing on current events. Stories on topics such as sports, entertainment news, and stock market reports were excluded from the data set. The following table is a summary of the files included in this release.</p><br> <table><br> <tbody><br> <tr><br> <td width="147">Source</td><br> <td width="125">Program</td><br> <td width="162">Epoch (YYYY.MM)</td><br> <td width="96">#hours</td><br> <td width="105">#words</td><br> </tr><br> <tr><br> <td>Dubai TV</td><br> <td>Dubai News</td><br> <td>2005.02 - 2005.12</td><br> <td>2.0</td><br> <td>11,078</td><br> </tr><br> <tr><br> <td>Nile TV</td><br> <td>News</td><br> <td>2000.11 - 2000.12</td><br> <td>0.5</td><br> <td>3,079</td><br> </tr><br> <tr><br> <td>Radio Sawa</td><br> <td>News at 06:00</td><br> <td>2005.11</td><br> <td>2.8</td><br> <td>9,712</td><br> </tr><br> <tr><br> <td>Voice of America</td><br> <td>News</td><br> <td>2000.11 - 2001.03</td><br> <td>5.4</td><br> <td>32,305</td><br> </tr><br> </tbody><br> </table><br> <h3>Transcription</h3><br> <p>The selected audio files were carefully transcribed by LDC annotators and professional transcription agencies following LDC's Quick Rich Transcription specification. Manual sentence units/segments (SU) annotation was also performed as part of the transcription task. Three types of end of sentence SU are identified:</p><br> <p>- statement SU - question SU - incomplete SU</p><br> <h3>Translation</h3><br> <p>After transcription and SU annotation, files were reformatted into a human-readable translation format and were assigned to professional translators for careful translation. Translators followed LDC's GALE Translation guidelines, which describe the makeup of the translation team, the source data format, the translation data format, best practices for translating certain linguistic features (such as names and speech disfluencies), and quality control procedures applied to completed translations.</p><br> <h3>Final Data</h3><br> <p>TDF Format</p><br> <p>All final data are in Tab Delimited Format (TDF). TDF is compatible with other transcription formats, such as Transcriber format and AG format, and it is easy to process.</p><br> <p>Each line of a TDF file corresponds to a speech segment and contains 13 tab delimited fields (the 13th field "suType" might be empty):</p><br> <p>field data_type ----- --------- 1 file unicode 2 channel int 3 start float 4 end float 5 speaker unicode 6 speakerType unicode 7 speakerDialect unicode 8 transcript unicode 9 section int 10 turn int 11 segment int 12 sectionType unicode 13 suType unicode</p><br> <p>A source TDF file and its translation are the same except that the transcript in the source TDF is replaced by its English translation.</p><br> <p>Encoding</p><br> <p>All data are encoded in UTF8.</p><br> <h3>Sponsorship</h3><br> <p>This work was supported in part by the Defense Advanced Research Projects Agency, GALE Program Grant No. HR0011-06-1-0003. The content of this publication does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.</p><br> <h3>Samples</h3><br> <p>For an example of the data in this corpus, please examine these screen captures of the <a href="desc/addenda/LDC2008T09_s.gif" rel="nofollow">source</a> and <a href="desc/addenda/LDC2008T09_t.gif" rel="nofollow">translation</a>.</p></br> Portions © 2005 Dubai TV, © 2000 Nile TV, © 2005-2007, 2008 Trustees of the University of Pennsylvania
<h3>引言</h3><br> <p>GALE第一阶段阿拉伯语广播新闻平行语料库(第二部分)是三卷本GALE第一阶段阿拉伯语广播新闻平行语料库的第二分册,该语料库与其他语料一同作为美国国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)资助的GALE项目第一年(第一阶段)的训练数据。<a href="http://catalog.ldc.upenn.edu/LDC2007T24" rel="nofollow">GALE第一阶段阿拉伯语广播新闻平行语料库(第一部分)</a>于2007年发布。</p><br> <p>GALE第一阶段阿拉伯语广播新闻平行语料库(第二部分)包含从多渠道采集的10.7小时阿拉伯语广播新闻节目的转写文本与英文译文。本语料库不包含用于生成转写文本及译文的原始音频文件。</p><br> <p>语言数据联盟(Linguistic Data Consortium, LDC)已发布以下GALE第一阶段与第二阶段阿拉伯语平行语料数据集:</p><br> <ul><br> <li>GALE第一阶段阿拉伯语广播新闻平行语料库(第一部分)(<a href="../../../LDC2007T24">LDC2007T24</a>)</li><br> <li>GALE第一阶段阿拉伯语广播新闻平行语料库(第二部分)(<a href="../../../LDC2008T09">LDC2008T09</a>)</li><br> <li>GALE第一阶段阿拉伯语博客平行语料库(<a href="../../../LDC2008T02">LDC2008T02</a>)</li><br> <li>GALE第一阶段阿拉伯语新闻组平行语料库(第一部分)(<a href="../../../LDC2009T03">LDC2009T03</a>)</li><br> <li>GALE第一阶段阿拉伯语新闻组平行语料库(第二部分)(<a href="../../../LDC2009T09">LDC2009T09</a>)</li><br> <li>GALE第二阶段阿拉伯语广播访谈平行语料库(第一部分)(<a href="../../../LDC2012T06">LDC2012T06</a>)</li><br> <li>GALE第二阶段阿拉伯语广播访谈平行语料库(第二部分)(<a href="../../../LDC2012T14">LDC2012T14</a>)</li><br> <li>GALE第二阶段阿拉伯语新闻专线平行语料库(<a href="../../../LDC2012T17">LDC2012T17</a>)</li><br> <li>GALE第二阶段阿拉伯语广播新闻平行语料库(<a href="../../../LDC2012T18">LDC2012T18</a>)</li><br> <li>GALE第二阶段阿拉伯语网络平行语料库(<a href="../../../LDC2013T01">LDC2013T01</a>)</li><br> </ul><br> <h3>源数据</h3><br> <p>本次采集的10.7小时阿拉伯语广播新闻录音选自4个来源与4档不同节目。</p><br> <p>本次数据采用人工遴选流程,选取适配GALE项目需求的内容,即聚焦时事的新闻与访谈节目。涉及体育、娱乐新闻、股市报道等主题的内容均被排除在本数据集之外。下表为本批次发布文件的汇总信息。</p><br> <table><br> <tbody><br> <tr><br> <td width="147">来源</td><br> <td width="125">节目名称</td><br> <td width="162">时间区间(YYYY.MM)</td><br> <td width="96">时长/小时</td><br> <td width="105">词汇量</td><br> </tr><br> <tr><br> <td>迪拜电视台(Dubai TV)</td><br> <td>迪拜新闻(Dubai News)</td><br> <td>2005.02 - 2005.12</td><br> <td>2.0</td><br> <td>11,078</td><br> </tr><br> <tr><br> <td>尼罗河电视台(Nile TV)</td><br> <td>新闻(News)</td><br> <td>2000.11 - 2000.12</td><br> <td>0.5</td><br> <td>3,079</td><br> </tr><br> <tr><br> <td>萨瓦广播电台(Radio Sawa)</td><br> <td>早间6点新闻(News at 06:00)</td><br> <td>2005.11</td><br> <td>2.8</td><br> <td>9,712</td><br> </tr><br> <tr><br> <td>美国之音(Voice of America)</td><br> <td>新闻(News)</td><br> <td>2000.11 - 2001.03</td><br> <td>5.4</td><br> <td>32,305</td><br> </tr><br> </tbody><br> </table><br> <h3>转写处理</h3><br> <p>经遴选的音频文件由LDC标注人员与专业转写机构按照LDC的快速富转写(Quick Rich Transcription)规范进行严谨转写。转写任务同时包含人工语句单元(Speech Unit, SU)标注环节。本次共识别三类语句单元结束类型:</p><br> <p>- 陈述型语句单元 - 疑问型语句单元 - 不完整语句单元</p><br> <h3>译制流程</h3><br> <p>完成转写与语句单元标注后,文件将被重格式化为便于人工阅读的译制格式,并交由专业译员进行精准翻译。译员需遵循LDC发布的GALE译制指南,该指南涵盖译制团队构成、源数据格式、译制数据格式、特定语言特征(如专有名词与言语不流畅现象)的译制规范,以及针对完成译稿的质量管控流程。</p><br> <h3>最终数据</h3><br> <p>TDF格式</p><br> <p>所有最终数据均采用制表符分隔格式(Tab Delimited Format, TDF)。该格式兼容Transcriber格式、AG格式等其他转写格式,且易于处理。</p><br> <p>TDF文件的每一行对应一段语音片段,包含13个制表符分隔的字段(第13个字段“suType”可为空):</p><br> <p>字段序号 字段名称 数据类型<br>1 文件 统一码(Unicode)<br>2 声道 整数(int)<br>3 起始时间 浮点数(float)<br>4 结束时间 浮点数(float)<br>5 说话人 统一码(Unicode)<br>6 说话人类型 统一码(Unicode)<br>7 说话人口音 统一码(Unicode)<br>8 转写文本 统一码(Unicode)<br>9 章节 整数(int)<br>10 话轮 整数(int)<br>11 片段 整数(int)<br>12 章节类型 统一码(Unicode)<br>13 语句单元类型(suType) 统一码(Unicode)</p><br> <p>源语言TDF文件与对应的译制文件结构一致,仅将源语言TDF中的转写文本替换为英文译文。</p><br> <p>编码方式</p><br> <p>所有数据均采用UTF8编码格式。</p><br> <h3>资助说明</h3><br> <p>本项目部分由美国国防高级研究计划局(Defense Advanced Research Projects Agency)通过GALE项目资助(资助编号HR0011-06-1-0003)。本出版物内容未必代表美国政府的立场或政策,不应被视为获得官方背书。</p><br> <h3>数据样例</h3><br> <p>若需查看本语料库的数据示例,请参阅<a href="desc/addenda/LDC2008T09_s.gif" rel="nofollow">源文件</a>与<a href="desc/addenda/LDC2008T09_t.gif" rel="nofollow">译制文件</a>的屏幕截图。</p><br> 部分内容 © 2005 迪拜电视台、© 2000 尼罗河电视台、© 2005-2007、2008 宾夕法尼亚大学托管会



