survey-language-technologies
收藏资源简介:
该数据集名为“社会经济地位如何影响语言技术交互的AI差距”,包含了来自不同社会经济背景的1000个个体的回应。这些数据用于研究社会经济地位对与生成式AI和大型语言模型交互的影响。参与者提供了人口统计和社会经济数据,以及他们之前向LLM提交的最多10个真实提示。
The dataset, titled "How Socioeconomic Status Shapes the AI Gap in Language Technology Interactions", contains responses from 1,000 individuals across diverse socioeconomic backgrounds. These data are employed to investigate the impact of socioeconomic status on interactions with generative AI and large language models (LLMs). Participants provided demographic and socioeconomic data, as well as up to 10 real prompts they had previously submitted to LLMs.
数据集概述
基本信息
- 名称: The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
- 许可证: MIT
- 语言: 英语 (en)
- 数据规模: 1K<n<10K
数据集摘要
- 样本量: 1,000名来自不同社会经济背景的个体
- 数据内容:
- 参与者提供的 demographic 和 socioeconomic 数据
- 参与者提交给大型语言模型(如ChatGPT)的真实 prompts,总计 6,482 条 unique prompts
- 研究目的: 研究社会经济地位(SES)如何影响与语言技术(特别是生成式AI和大型语言模型)的互动
数据集结构
- 文件格式: CSV (
survey_language_technologies.csv) - 字段说明:
- 人口统计信息:
gender,age,nationality,ethnicity,marital,language,religion - 教育背景:
education,mum_education,dad_education - 社会经济状况:
ses,home,employment,occupation,mother_occupation,father_occupation - 兴趣爱好与技术使用:
hobbies,tech,know_nlp,use_nlp,would_nlp - LLM使用情况:
frequency_llm,llm_use,usecases,contexts - 用户提交的prompts:
prompt1–prompt10 - 其他:
comments
- 人口统计信息:
引用信息
bibtex @inproceedings{bassignana-2025-survey, title = "The {AI} {G}ap: {H}ow {S}ocioeconomic {S}tatus {A}ffects {L}anguage {T}echnology {I}nteractions", author = "Bassignana, Elisa and Cercas Curry, Amanda and Hovy, Dirk", booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics", year = "2025", url = "https://arxiv.org/abs/2505.12158" }
数据集维护者
- Elisa Bassignana (IT University of Copenhagen)
- Amanda Cercas Curry (CENTAI Institute)
- Dirk Hovy (Bocconi University)
相关链接
- 数据集文件:
survey_language_technologies.csv - 调查界面: https://nlp-use-survey.streamlit.app/
- 论文预印本: https://arxiv.org/abs/2505.12158




