HEBDB
收藏资源简介:
HEBDB是由耶路撒冷希伯来大学和以色列理工学院共同创建的一个用于希伯来语语音处理的数据集。该数据集包含约2500小时的自然和即兴希伯来语语音记录,涵盖了多种说话者和主题。数据集的创建旨在推动希伯来语语音处理工具的研究和开发,特别是在自动语音识别(ASR)领域。数据集包括原始录音和预处理版本,预处理版本经过语音活动检测和自动转录,更适合用于训练声学模型。HEBDB的应用领域主要集中在人工智能和语音技术,旨在解决低资源语言在语音处理方面的挑战。
HEBDB is a dataset for Hebrew speech processing co-created by the Hebrew University of Jerusalem and the Technion – Israel Institute of Technology. It contains approximately 2,500 hours of natural and spontaneous Hebrew speech recordings, covering a wide range of speakers and topics. The dataset was developed to advance research and development of Hebrew speech processing tools, particularly in the field of automatic speech recognition (ASR). It includes both raw audio recordings and preprocessed versions. The preprocessed variants have undergone voice activity detection and automatic transcription, making them more suitable for acoustic model training. HEBDB is primarily applied in artificial intelligence and speech technology, aiming to address the challenges in speech processing for low-resource languages.
HebDB 数据集概述
数据集名称
HebDB
数据集描述
HebDB 是一个用于希伯来语语音处理的弱监督数据集。该数据集提供了大约 2500 小时的自然和即兴的希伯来语语音记录,包含多种说话者和话题。
关键词
HebDB, 深度学习, 语音, 数据集, 希伯来语语音
作者
- Arnon Turetzky<sup>1</sup>
- Or Tal<sup>1</sup>
- Yael Segal<sup>2</sup>
- Yehoshua Dissen<sup>2</sup>
- Ella Zeldes<sup>1</sup>
- Amit Roth<sup>1</sup>
- Eyal Cohen<sup>2</sup>
- Yosi Shrem<sup>2</sup>
- Roni Chernyak<sup>2</sup>
- Olga Seleznova<sup>2</sup>
- Joseph Keshet<sup>2</sup>
- Yossi Adi<sup>1</sup>
机构
- <sup>1</sup>耶路撒冷希伯来大学
- <sup>2</sup>以色列理工学院

- 1HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing耶路撒冷希伯来大学, 以色列理工学院 · 2024年



