STORIUM
收藏资源简介:
STORIUM数据集是由麻省大学阿默斯特分校与STORIUM合作创建的,旨在推动机器参与的故事生成研究。该数据集包含5743个长篇故事,总计125M个tokens,每个故事都包含精细的自然语言注释,如角色目标和属性,这些注释分布在每个叙述中,为指导模型提供了坚实的依据。STORIUM数据集的特点在于其故事的长度和丰富性,以及通过STORIUM平台收集的真实用户数据。数据集的应用领域主要集中在故事生成模型的训练和评估,特别是在机器参与的故事创作过程中,通过让真实作者与模型互动,评估模型的生成效果。
The STORIUM Dataset was co-developed by the University of Massachusetts Amherst and STORIUM, with the goal of advancing research on machine-involved story generation. This dataset comprises 5,743 full-length stories, totaling 125 million tokens. Each story is equipped with fine-grained natural language annotations such as character goals and attributes, which are integrated throughout the narrative, providing solid grounding for guiding models. What distinguishes the STORIUM Dataset is its lengthy, rich story corpus and real user data collected via the STORIUM platform. Its main application scenarios focus on the training and evaluation of story generation models, particularly in evaluating the generation quality of models during human-machine collaborative story creation, where real authors interact with the models.




