Zema Dataset
收藏资源简介:
Zema Dataset是由中央研究院信息科学研究所等机构共同创建的一个专门用于分析埃塞俄比亚东正教特瓦赫多教会(EOTC)圣歌的数据集。该数据集包含10小时的音频数据,共369条实例,涵盖了详细的单词级别时间边界、阅读音调标注以及圣歌模式标签。数据来源于Eat the Book网站,经过严格的预处理和质量控制,包括音频清理、分段和文本提取。该数据集旨在支持圣歌模式分类、歌词转录、歌词与音频对齐以及音乐生成等任务,推动对埃塞俄比亚独特宗教音乐文化的保护与研究。
The Zema Dataset is a specialized dataset co-developed by the Institute of Information Science, Academia Sinica, and other institutions for the analysis of chant music of the Ethiopian Orthodox Tewahedo Church (EOTC). It contains 10 hours of audio data and a total of 369 instances, which include detailed word-level temporal boundaries, recitation tone annotations, and chant mode labels. Sourced from the Eat the Book website, the dataset has undergone rigorous preprocessing and quality control procedures, including audio cleaning, segmentation, and text extraction. This dataset aims to support tasks such as chant mode classification, lyric transcription, lyric-audio alignment, and music generation, thereby promoting the preservation and academic research of Ethiopia's unique religious musical culture.

- 1Zema Dataset: A Comprehensive Study of Yaredawi Zema with a Focus on Horologium Chants中央研究院信息科学研究所, 台湾大学信息管理系统, 台湾科技大学信息管理系, 台湾科技大学材料科学与工程系, 台湾国际研究生院社交网络与人类中心计算项目, 台湾政治大学计算机科学系, 贡德尔大学信息系统系 · 2024年



