遇见数据集

FM2 (FoolMeTwice)

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

FoolMeTwice(简称 FM2)是通过有趣的多人游戏收集的具有挑战性的蕴涵对的大型数据集。游戏化鼓励对抗性示例,与其他流行的蕴涵数据集相比,大大降低了可以使用“捷径”解决的示例数量。玩家有两个任务。第一个任务要求玩家根据来自维基百科页面的证据写一个合理的声明。第二个显示了其他玩家写的两个似是而非的说法,其中一个是错误的,目标是在时间用完之前识别它。玩家“付费”以查看从证据池中检索到的线索:玩家需要的证据越多,索赔越难。有动机的玩家之间的博弈导致制定声明的不同策略,例如时间推断和转移到不相关的证据,并为蕴涵和证据检索任务带来更高质量的数据。

FoolMeTwice (abbreviated as FM2) is a large-scale dataset of challenging entailment pairs collected via an engaging multiplayer game. The game's mechanics encourage adversarial examples, which drastically reduce the number of examples that can be solved via "shortcuts" compared to other popular entailment datasets. Players are assigned two tasks. The first task requires players to compose a plausible claim using evidence sourced from Wikipedia pages. The second task presents two plausible claims written by other players, one of which is false, and the goal is to identify the false claim before the time limit runs out. Players "pay" to access clues retrieved from a pool of evidence: the more evidence a player needs, the more difficult the corresponding claim becomes. The strategic interactions between motivated players foster diverse claim-generation strategies, such as temporal inference and shifting to irrelevant evidence, ultimately yielding higher-quality data for both entailment and evidence retrieval tasks.

提供机构:
OpenDataLab
创建时间:
2022-05-23
搜集汇总
数据集介绍
FM2 (FoolMeTwice) 数据集图片
背景与挑战
背景概述
FM2(FoolMeTwice)是一个通过多人游戏收集的大型蕴涵对数据集,其游戏化设计鼓励对抗性示例,显著减少了可被“捷径”解决的样本。玩家通过撰写声明和识别错误声明的任务,结合证据检索机制,生成了高质量的自然语言推理数据。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务