Uli Dataset
收藏资源简介:
Uli数据集是由哥本哈根大学等机构创建,专注于性别暴力在线检测的数据集。该数据集包含三种语言:印地语、泰米尔语和印度英语,共计约8000条推文,由专家标注性别虐待相关内容。数据集的创建过程采用了参与式方法,旨在通过AI系统解决性别暴力问题。此数据集不仅关注文本内容,还考虑了性别虐待的上下文和经验,为性别暴力在线检测提供了丰富的资源。
The Uli Dataset is a collection developed by institutions including the University of Copenhagen, dedicated to online detection of gender-based violence. This dataset covers three languages: Hindi, Tamil, and Indian English, containing approximately 8,000 tweets annotated by experts for gender-based abusive content. The dataset was constructed using a participatory methodology, aiming to address gender-based violence through AI systems. Beyond focusing on textual content, this dataset also takes into account the context and lived experiences related to gender-based abuse, providing a rich resource for online gender-based violence detection.

- 1The Uli Dataset: An Exercise in Experience Led Annotation of oGBV哥本哈根大学, 丹麦 · 2023年



