Who are you,ChatGPT?Personality and Demographic Style in LLM-Generated Content
收藏资源简介:
本数据集收集了来自Reddit的开放式问题及其答案,以及LLMs对这些问题生成的回复,旨在研究LLMs的语言是否表现出与人类相似的个性和人口统计特征。数据集包含来自175个Reddit社区的13K个帖子以及超过30K条评论,由数千名Reddit用户撰写。此外,数据集还包含了来自多个LLMs的回复,用于比较人类和模型在Big Five维度上的个性和人口统计特征。该数据集可用于研究LLMs的个性和人口统计特征,以及它们在自然语言处理中的应用。
This dataset collects open-ended questions and their corresponding human answers from Reddit, as well as responses generated by large language models (LLMs), aiming to investigate whether the language of LLMs exhibits personality and demographic characteristics similar to those of humans. The dataset includes 13,000 posts and over 30,000 comments from 175 Reddit communities, authored by thousands of Reddit users. Additionally, the dataset contains responses from multiple LLMs, which are used to compare the personality and demographic characteristics of humans and models across the Big Five personality dimensions. This dataset can be utilized to study the personality and demographic traits of LLMs, as well as their applications in natural language processing.




