遇见数据集

AIML-TUDA/Rail2Country

收藏
Hugging Face2026-05-21 更新2026-01-03 收录
官方服务:

资源简介:

这是与论文《ActivationReasoning: Latent Activation Spaces中的逻辑推理》对应的数据集。Rail2Country数据集是在该论文中引入的一个新颖基准,专门设计用于测试大型语言模型(LLMs)在概念以间接或抽象方式表达时进行演绎推理的能力。该数据集受早期基于火车的推理任务启发,每个实例要求根据编码到逻辑组件中的一组国旗颜色到国家的映射规则,将火车的车厢颜色序列映射到其原产国。数据集包括两个不同的变体以评估泛化能力:* R2C-Mono:在此设置中,颜色在火车描述中明确陈述(例如“红色”)。* R2C-Meta:此变体用比喻和描述性代理替换明确的颜色提及(例如“像番茄一样颜色”),迫使模型通过整合上下文措辞和世界知识来解决隐式线索。该数据集使研究人员能够调查LLMs从显式词汇概念泛化到抽象、元级描述的能力。

This is the dataset corresponding to the paper "ActivationReasoning: Logical Reasoning in Latent Activation Spaces". The Rail2Country dataset is a novel benchmark introduced in the paper "ActivationReasoning: Logical Reasoning in Latent Activation Spaces". It is specifically designed to test the ability of Large Language Models (LLMs) to perform deductive reasoning over concepts when those concepts are expressed indirectly or abstractly. Inspired by earlier train-based reasoning tasks, each instance requires mapping a trains car color sequence to its country of origin, based on a set of flag color-to-country rules encoded into the logic component. The dataset includes two distinct variants to evaluate generalization: * R2C-Mono: In this setting, colors are explicitly stated in the train descriptions (e.g., "red"). * R2C-Meta: This variant replaces explicit color mentions with similes and descriptive proxies (e.g., "colored like a tomato"), forcing the model to resolve implicit cues by integrating contextual phrasing and world knowledge. This dataset enables researchers to investigate how well LLMs can generalize from explicit lexical concepts to abstract, meta-level descriptions.

提供机构:
AIML-TUDA
二维码
社区交流群
二维码
科研交流群
商业服务