大型日本网络语料库
收藏资源简介:
大型日本网络语料库是由东京工业大学计算机科学与技术学院创建的,旨在为大型语言模型提供高质量的日语训练数据。该数据集包含约3121亿字符,覆盖了2020至2023年间爬取的约634亿网页中的17300万页,是所有可用日语训练语料库中最大的。创建过程中,研究团队从Common Crawl档案中提取并精炼文本,特别设计了针对日语文本的过滤方法,以确保数据质量。该数据集主要用于训练日语大型语言模型,解决日语处理中的性能问题,提升模型在日语基准数据集上的表现。
The Large Japanese Web Corpus was developed by the School of Computer Science and Technology, Tokyo Institute of Technology, with the core objective of providing high-quality Japanese training data for large language models (LLMs). Comprising approximately 173 million pages extracted from around 63.4 billion web pages crawled between 2020 and 2023, the corpus has a total size of about 312.1 billion characters, making it the largest available Japanese training corpus to date. During its development, the research team extracted and refined textual content from Common Crawl archives, and designed specialized filtering methods tailored specifically for Japanese text to ensure data quality. This corpus is primarily used for training Japanese large language models, aiming to address performance issues in Japanese language processing and improve the models' performance on Japanese benchmark datasets.

- 1Building a Large Japanese Web Corpus for Large Language Models东京工业大学计算机科学与技术学院 · 2024年



