遇见数据集

CSC Deceptive Speech

收藏
DataCite Commons2023-01-11 更新2024-07-13 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>CSC Deceptive Speech was developed by Columbia University, SRI International and University of Colorado Boulder. It consists of 32 hours of audio interviews from 32 native speakers of Standard American English (16 male,16 female) recruited from the Columbia University student population and the community. The purpose of the study was to distinguish deceptive speech from non-deceptive speech using machine learning techniques on extracted features from the corpus.</p><br> <p>The participants were told that they were participating in a communication experiment which sought to identify people who fit the profile of the top entrepreneurs in America. To this end, the participants performed tasks and answered questions in six areas. They were later told that they had received low scores in some of those areas and did not fit the profile. The subjects then participated in an interview where they were told to convince the interviewer that they had actually achieved high scores in all areas and that they did indeed fit the profile. The task of the interviewer was to determine how he thought the subjects had actually performed, and he was allowed to ask them any questions other than those that were part of the performed tasks. For each question from the interviewer, subjects were asked to indicate whether the reply was true or contained any false information by pressing one of two pedals hidden from the interviewer under a table.</p><br> <h3>Data</h3><br> <p>Interviews were conducted in a double-walled sound booth and recorded to digital audio tape on two channels using Crown CM311A Differoid headworn close-talking microphones, then downsampled to 16kHz before processing.</p><br> <p>The interviews were orthographically transcribed by hand using the NIST EARS transcription guidelines. Labels for local lies were obtained automatically from the pedal-press data and hand-corrected for alignment, and labels for global lies were annotated during transcription based on the known scores of the subjects versus their reported scores. The orthographic transcription was force-aligned using the SRI telephone speech recognizer adapted for full-bandwidth recordings. There are several segmentations associated with the corpus: the implicit segmentation of the pedal presses, derived semi-automatically sentence-like units (EARS SLASH-UNITS or SUs) which were hand labeled, intonational phrase units and the units corresponding to each topic of the interview.</p><br> <p>Transcript files are in .trs format and audio files are .wav presented in <a href="https://xiph.org/flac/" rel="nofollow">flac-compressed</a> form for this release.</p><br> <h3>Samples</h3><br> <p>Please view these <a href="desc/addenda/LDC2013S09.wav" rel="nofollow">audio</a> and <a href="desc/addenda/LDC2013S09.txt" rel="nofollow">transcript</a> samples for the interviewer side of a conversation..</p><br> <h3>Updates</h3><br> <p>On May 22, 2014 an additional documentation file was added to explain the questions&nbsp; participants were asked.</p></br> Portions © 2013 The Trustees of Columbia University, Trustees of the University of Pennsylvania

### 引言 CSC欺骗性言语数据集(CSC Deceptive Speech)由哥伦比亚大学、SRI国际研究院及科罗拉多大学博尔德分校联合研发。该数据集包含32名以标准美式英语为母语的受访者的32小时音频访谈内容,其中男性、女性各16名,招募渠道涵盖哥伦比亚大学在校生及本地社区人群。本研究旨在通过对语料库提取的特征应用机器学习技术,实现欺骗性言语与非欺骗性言语的自动区分。 受访参与者被告知,他们正在参与一项旨在筛选符合美国顶尖企业家特质人群的沟通实验。为此,参与者需完成六项任务并回答对应领域的问题。随后参与者被告知,他们在部分任务中得分较低,并不符合顶尖企业家的特质画像。此后受试者参与一场访谈,被要求说服访谈者,称自己在所有任务中均取得高分,且确实符合该特质画像。访谈者的任务是判断受试者实际的任务完成情况,且除任务本身包含的问题外,可向受试者提出任意问题。针对访谈者提出的每个问题,受试者需通过按压桌下隐藏的两个踏板之一,来表明自己的回答是否属实或包含虚假信息,该操作对访谈者不可见。 ### 数据说明 访谈在双层隔音隔声间内进行,采用Crown CM311A Differoid头戴式近距麦克风双声道录制至数字音频磁带,后续处理前将采样率降为16kHz。 访谈内容按照NIST EARS转录规范进行人工正字法转录。局部谎言标签可从踏板按压数据中自动提取,并经人工校正以实现对齐;全局谎言标签则在转录过程中,基于受试者实际得分与自述得分的对比进行标注。正字法转录结果通过适配全带宽录音的SRI电话语音识别器进行强制对齐。该语料库包含多种分割方式:踏板按压事件的隐式分割、半自动生成的类句单元(EARS SLASH-UNITS,简称SUs,经人工标注)、语调短语单元,以及对应访谈每个话题的单元。 本版本发布的转录文件格式为.trs,音频文件则为经FLAC压缩的.wav格式。 ### 样本示例 请查看以下对话访谈者端的<a href="desc/addenda/LDC2013S09.wav" rel="nofollow">音频样本</a>与<a href="desc/addenda/LDC2013S09.txt" rel="nofollow">转录文本样本</a>。 ### 更新记录 2014年5月22日,新增一份说明文档,用于解释本次实验向参与者提出的问题。 部分内容 © 2013 哥伦比亚大学理事会、宾夕法尼亚大学理事会 版权所有。

创建时间:
2020-11-30
搜集汇总
数据集介绍
CSC Deceptive Speech 数据集图片
背景与挑战
背景概述
CSC Deceptive Speech是一个英语语音数据集,包含32小时的音频访谈,来自32名标准美式英语母语者,旨在研究欺骗性语音的识别。数据通过实验设计收集,参与者在访谈中标记回答的真实性,音频以16kHz采样率录制,并附有手动转录和自动标注的欺骗标签。该数据集适用于语音识别和异常分析等机器学习应用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务