Japanese Corpus for Human-AI Talks (J-CHAT)
收藏资源简介:
日本人类-AI对话语料库(J-CHAT)是由东京大学和庆应大学联合创建的大规模日语口语对话数据集。该数据集包含69,000小时的语音数据,来源于YouTube和播客,旨在提供自然且清晰的对话样本。数据集的创建过程包括自动化的数据收集、语言识别、对话提取和噪音去除。J-CHAT主要用于训练对话导向的口语语言模型,以提高人机交互的自然性和有效性。
The Japanese Human-AI Dialogue Corpus (J-CHAT) is a large-scale spoken Japanese dialogue dataset jointly developed by the University of Tokyo and Keio University. This corpus contains 69,000 hours of speech data sourced from YouTube and podcasts, and is designed to provide natural and clear conversational samples. The dataset construction process includes automated data collection, speech recognition, dialogue extraction and noise removal. J-CHAT is primarily used for training dialogue-oriented spoken language models, aiming to enhance the naturality and effectiveness of human-computer interaction.

- 1J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling东京大学, 庆应大学 · 2024年



