This is the stand-off GrAF version of Spanish portions of the Wikipedia (based on a 2006 dump). This Wikipedia Spanish Corpus contains 257019 articles that contain about 150,1 million words in raw tex
该数据集包含两个版本的Estonian National Corpus 2021,分别是有形态学标注的文本(corpus_et.jsonl)和清理后的纯文本(corpus_et_clean.jsonl)。数据集总大小约为43GB,包含约1.96亿个句子、24亿个单词、1170万篇文档和6450万个段落。这些数据可以用于形态学分析、自然语言理解、语言模型微调等多种自然语言处理任务。