Novelist-CoT
收藏资源简介:
Novelist-CoT 是一个专为监督式微调和风格化叙事生成设计的长篇创意写作数据集。该数据集整合了多种生成流程,采用统一结构,并在接受的样本中保留了规划痕迹(`<think>`)和最终散文(`<answer>`)。数据集强调长篇故事的延续与扩展、风格条件写作、叙事规划质量、多语言翻译变体以及用于下游过滤的结构化子类型标记。数据集包含 8,163 行数据,总标记数为 38,322,752(cl100k_base),平均每行 4,694.69 个标记。每行包含一个 `type` 字段,便于过滤为专门的训练切片。数据集支持多种语言翻译,包括阿拉伯语、中文、法语、德语等。所有记录采用统一的 JSON 模式,包含 `hash`、`type`、`instruction`、`input`、`output` 和 `metadata` 等字段。数据集通过合并多个版本的文件构建,并经过严格的质量筛选,确保长篇叙事的连贯性。推荐用于叙事生成的监督微调、延续和章节扩展训练、思维链感知的写作系统等应用场景。
Novelist-CoT is a long-form creative writing dataset designed specifically for supervised fine-tuning and stylized narrative generation. This dataset integrates multiple generation workflows, adopts a unified structure, and retains planning traces (`<think>`) and final prose (`<answer>`) in the accepted samples. The dataset emphasizes the continuation and expansion of long-form stories, style-conditioned writing, narrative planning quality, multilingual translation variants, and structured subtype tags for downstream filtering. It contains 8,163 rows of data, with a total token count of 38,322,752 (using the cl100k_base tokenizer), averaging 4,694.69 tokens per row. Each row includes a `type` field to facilitate filtering into specialized training slices. The dataset supports multilingual translations, including Arabic, Chinese, French, German, and other languages. All records adhere to a unified JSON schema, containing fields such as `hash`, `type`, `instruction`, `input`, `output`, and `metadata`. The dataset is constructed by merging multiple versions of files and subjected to rigorous quality filtering to ensure the coherence of long-form narratives. It is recommended for application scenarios including supervised fine-tuning for narrative generation, continuation and chapter expansion training, and chain-of-thought aware writing systems.



