strombergnlp/zulu_stance
收藏资源简介:
这是一个用于祖鲁语(Zulu)的立场检测数据集。数据由祖鲁语母语者从英语源文本翻译而来。该数据集的目标是促进祖鲁语中的立场检测,并衡量翻译中的领域转移。数据集包含1343个句子,数据字段包括id、text、target和stance。数据集的创建基于Semeval2016任务6的数据,并通过手动翻译生成。数据集的使用存在一定的社会影响和偏见问题,因为源文本来自英语用户,可能包含美国特定的主题和偏见。
This is a stance detection dataset for the Zulu language. The dataset was translated from English source texts into Zulu by native Zulu speakers. Its objective is to facilitate stance detection research in Zulu and measure domain shift resulting from cross-lingual translation. The dataset comprises 1,343 sentences, with data fields including id, text, target, and stance. It was developed based on the data from SemEval 2016 Task 6 and generated through manual translation. There are notable social impact and bias concerns associated with the usage of this dataset, as the source texts originate from English-speaking users and may contain US-specific topics and inherent biases.
数据集概述
数据集名称
- 名称: ZUstance
- 别名: zulu-stance
数据集属性
- 语言: Zulu (
bcp47:zu) - 许可证: CC-BY-4.0
- 多语言性: 单语种
- 大小: 1K<n<10K
- 来源: 原始数据
- 任务类别: 文本分类
- 任务ID: 事实核查, 情感分类
- 标签: 立场检测
数据集描述
- 概述: 这是一个Zulu语言的立场检测数据集。数据由Zulu母语者从英语源文本翻译而来。
- 目的: 旨在利用英语的进展,将知识转移到其他语言,特别是Zulu语言,通过领域适应技术减少域间差距。
数据集结构
- 数据实例: 示例包括ID、文本、目标和立场。
- 数据字段: 包括ID、文本、目标和立场,其中立场标签包括“FAVOR”, “AGAINST”, “NONE”。
- 数据分割: 训练集包含1343个句子。
数据集创建
- 采集与规范化: 原始数据来自Semeval2016任务6,后手动翻译为Zulu。
- 源语言生产者: 英语Twitter用户。
- 注释过程: 注释来自Semeval2016任务6。
使用数据注意事项
- 社会影响: 数据可能包含用户删除的内容,未经过滤,可能存在有害文本。
- 偏见讨论: 尽管数据为Zulu语言,但源文本来自英语Twitter用户,可能包含与Zulu社会不同的偏见和话题。
附加信息
- 数据集管理: 由论文作者管理。
- 许可信息: 根据CC-BY 4.0许可发布。
- 引用信息: 参考文献格式如README文件所示。




