新浪微博性别歧视审查(SWSR)数据集
收藏资源简介:
SWSR数据集是首个针对中文的性别歧视数据集,由伦敦玛丽女王大学创建。该数据集包含10496条新浪微博内容,包括微博及其评论,旨在识别和分析中文网络环境中的性别歧视言论。数据集通过关键词搜索收集,涵盖多种性别歧视类型,如外貌、文化背景、微侵犯和性侵犯。此外,数据集还提供用户性别和位置等匿名信息,以支持更深入的分析。SWSR数据集的应用领域包括性别歧视的自动检测和分析,以及促进跨语言性别歧视研究。
The SWSR dataset is the first Chinese-language dataset focused on gender discrimination, created by Queen Mary University of London. Comprising 10,496 Sina Weibo posts and their accompanying comments, this dataset is designed to identify and analyze gender-discriminatory remarks in Chinese online environments. Collected via keyword-based searches, the dataset covers multiple types of gender discrimination, including those targeting appearance, cultural background, microaggressions, and sexual assault. Additionally, the dataset provides anonymous user information such as gender and location to support more in-depth analyses. Applications of the SWSR dataset include automatic detection and analysis of gender discrimination, as well as advancing cross-linguistic gender discrimination research.




