threite/Bundestag-v2
收藏资源简介:
Bundestag-v2数据集是从ParlSpeech V2数据集中生成的,包含1990年至2020年德国议会的演讲,并标注了演讲者的政党。数据集的主要任务是文本分类,语言为德语。数据字段包括演讲的文本和演讲者的政党。数据集分为训练集、验证集和测试集。创建该数据集的目的是训练一个能够按政党分类演讲的语言模型。数据集的使用可能涉及社会影响,因为政治演讲内容可能具有争议性和潜在危害性。
The Bundestag-v2 dataset is derived from the ParlSpeech V2 dataset. It contains speeches from the German Bundestag between 1990 and 2020, annotated with the political party of each speaker. The primary task of this dataset is text classification, with all texts in German. The data fields include the speech text and the political party of the speaker. The dataset is split into training, validation, and test sets. The purpose of creating this dataset is to train a language model capable of classifying speeches based on the speakers' political parties. The usage of this dataset may involve social impacts, as political speech content can be controversial and potentially harmful.
数据集概述
数据集名称
- 名称: Bundestag-v2
- 别名: ParlSpeech V2
数据集基本信息
语言
- 语言: 德语
- 语言创建方式: 专家生成
许可
- 许可类型: CC0-1.0
多语言性
- 多语言性: 单语种
大小
- 数据集大小: 100K<n<1M
标签
- 标签: Bundestag, ParlSpeech
任务类别
- 任务类别: 文本分类
- 任务ID: 实体链接分类
数据集内容
数据集摘要
- 摘要: 该数据集包含1990年至2020年间德国议会的演讲,演讲者所属政党已标注。
支持的任务
- 任务: 文本分类
数据结构
- 数据字段:
- text: 德语演讲文本
- party: 演讲者所属政党
- 数据分割:
- 分割类型: 训练集, 验证集, 测试集
数据集创建
- 创建理由: 用于训练能够根据政党分类演讲的语言模型。
- 源数据: ParlSpeech V2
使用数据注意事项
- 社会影响: 由于包含政治演讲,内容可能具有争议性和潜在危害。
许可信息
- 许可: CCO 1.0
引用信息
-
引用格式:
@data{DVN/L4OAKN_2020, author = {Rauh, Christian and Schwalbach, Jan}, publisher = {Harvard Dataverse}, title = {{The ParlSpeech V2 data set: Full-text corpora of 6.3 million parliamentary speeches in the key legislative chambers of nine representative democracies}}, year = {2020}, version = {V1}, doi = {10.7910/DVN/L4OAKN}, url = {https://doi.org/10.7910/DVN/L4OAKN} }




