Reading with Intent Dataset
收藏资源简介:
该数据集名为Reading with Intent Dataset,由乔治亚理工学院的研究团队创建,旨在解决大语言模型在处理不同情感和语言风格文本时的挑战。数据集基于Natural Questions数据集,通过检索算法获取上下文段落,并使用多个大语言模型生成11种不同情感的语言风格,总计包含3,636,592条独特段落。数据集的创建过程包括从NQ数据集中检索上下文段落,并通过多个大语言模型生成不同情感的语言风格。该数据集的应用领域主要是情感翻译和阅读理解任务,旨在通过情感翻译模型将文本转换为中性语气,从而提升大语言模型在处理情感化文本时的表现。
The Reading with Intent Dataset was created by a research team at the Georgia Institute of Technology to address the challenges encountered by large language models (LLMs) when processing texts with diverse emotional tones and linguistic styles. Grounded in the Natural Questions dataset, this dataset acquires contextual paragraphs via retrieval algorithms, and leverages multiple large language models to generate linguistic styles aligned with 11 distinct emotions, resulting in a total of 3,636,592 unique paragraphs. The dataset's construction workflow involves retrieving contextual paragraphs from the Natural Questions dataset and generating emotionally varied linguistic styles through multiple LLMs. Its primary application domains are sentiment translation and reading comprehension tasks, with the objective of utilizing sentiment translation models to convert texts into neutral tones, thereby improving the performance of LLMs when handling emotionalized texts.




