多语言攻击性语言和仇恨言论检测数据集
收藏资源简介:
本研究开发了一个针对Hausa, Yoruba和Igbo三种尼日利亚主要语言的攻击性语言和仇恨言论检测的多语言数据集。数据集包含从Twitter收集并由母语者手动标注的推文,旨在通过自然语言处理技术自动检测和移除社交媒体中的攻击性和仇恨内容。数据集的创建考虑了语言的多样性和文化背景,适用于解决社交媒体中的语言攻击和仇恨言论问题。
This study develops a multilingual dataset for offensive language and hate speech detection targeting three dominant Nigerian languages: Hausa, Yoruba, and Igbo. The dataset comprises tweets collected from Twitter and manually annotated by native speakers, aiming to automatically detect and eliminate offensive and hateful content on social media via natural language processing (NLP) technologies. The construction of the dataset takes into account linguistic diversity and cultural context, making it suitable for resolving issues related to linguistic attacks and hate speech on social media platforms.

- 1A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages卡诺巴耶罗大学计算机科学系 · 2024年



