该数据集包含输入文本(input_text)和目标文本(target_text)两个字段,适用于研究文本生成任务。数据集分为训练集和验证集,其中训练集有10000个样本,验证集有1024个样本。数据集用于支持论文《Roll the dice & look before you leap: Going beyond the creative limits of next-token predicti
Indonesian text corpus from web. Crawling done by SpiderLing in 2017. Filtering by JusText and Onion (see http://corpus.tools/ for details). Tagged and lemmatized by MorphInd...