遇见数据集

UnQover

收藏
arXiv2025-09-30 收录
数据链接:
官方服务:

资源简介:

该数据集包含了一系列非负面问题,用于分析语言模型输出中的偏见。此外,值得注意的是,UnQover数据集并未提供“无法确定”的选项,但模型有时会返回“未知”的答案,这类答案被归类在“未知”类别中。这项任务旨在对跨语言模型进行偏见评估。

This dataset comprises a series of non-negative questions, designed for analyzing biases in the outputs of language models. It is worth noting that the UnQover dataset does not offer an "unable to determine" option, yet models occasionally return "unknown" answers, which are classified under the "unknown" category. This task is intended for bias evaluation of cross-lingual language models.

搜集汇总
数据集介绍
UnQover 数据集图片
背景与挑战
背景概述
UnQover是一个用于检测机器学习模型中刻板印象偏见的数据集和工具集,基于EMNLP 2020 Findings论文。它通过生成未指定问题来量化模型在性别、国籍、民族和宗教等维度的偏见,并提供完整的代码实现,包括数据生成、模型预测、分析和可视化功能,支持多种预训练模型和自定义训练流程。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务