DDSC/angry-tweets
收藏资源简介:
AngryTweets数据集由丹麦的匿名Twitter数据组成,这些数据通过众包进行了情感分析标注。数据集适用于情感分析任务,包含训练集和测试集,分别有2,437条和1,047条推文。每条数据包含推文内容和情感标签,标签分为positiv、neutral和negativ。数据集由Amalie Brogaard Pauli等人创建,并遵循CC BY 4.0许可证。
The AngryTweets dataset comprises anonymized Twitter data sourced from Denmark, which was annotated for sentiment analysis via crowdsourcing. This dataset is intended for sentiment analysis tasks, and includes a training set and a test set with 2,437 and 1,047 tweets respectively. Each entry contains the tweet content and a sentiment label, with the available labels being positive, neutral, and negative. This dataset was created by Amalie Brogaard Pauli et al. and is licensed under CC BY 4.0.
数据集概述
数据集基本信息
- 名称: AngryTweets
- 语言: 丹麦语 (da)
- 许可证: CC BY 4.0
- 多语言性: 单语种
- 大小: 1K<n<10K
- 来源: 原始数据
- 任务类别: 文本分类
- 任务ID: 情感分类
数据集描述
数据集总结
- 内容: 包含匿名的丹麦语Twitter数据,用于情感分析。
- 创建方式: 通过众包进行标注。
- 参考文献: Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021.
支持的任务和排行榜
- 任务: 情感分析
数据集结构
数据实例
- 结构: 每个实例包含一条推文及其情感标签。
数据字段
text(str): 推文内容。label(str): 情感标签,可以是 "positiv"(正面)、"neutral"(中性)或 "negativ"(负面)。
数据分割
- 分割方式: 训练集和测试集,测试集占30%,随机分层抽样。
- 数量: 训练集包含2,437条推文,测试集包含1,047条推文。
附加信息
数据集创建者
- 创建者: Amalie Brogaard Pauli, Maria Barrett, Ophélie Lacroix, Rasmus Hvingelby
- 推文匿名化: @saattrupdan
许可证信息
- 许可证: CC BY 4.0
引用信息
@inproceedings{pauli2021danlp, title={DaNLP: An open-source toolkit for Danish Natural Language Processing}, author={Pauli, Amalie Brogaard and Barrett, Maria and Lacroix, Oph{e}lie and Hvingelby, Rasmus}, booktitle={Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa)}, pages={460--466}, year={2021} }




