Multilingual Native Reasoning Challenge (MultiNRC)
收藏资源简介:
MultiNRC是一个包含超过1000个由法语、西班牙语和中文母语者编写的本族语言和文化背景下的推理问题的评估基准。数据集涵盖了语言特定的语言推理、文字游戏和谜语、文化/传统推理以及与文化相关的数学推理四个核心推理类别。数据集的创建过程包括招募母语者创作具有挑战性的推理问题,并提供客观和简短的最终答案,以方便自动评估。MultiNRC旨在解决当前大型语言模型在多语言推理能力方面的不足,并促进多语言和具有文化背景的评估研究。
MultiNRC is an evaluation benchmark encompassing over 1,000 reasoning questions developed by native speakers of French, Spanish, and Chinese, grounded in their respective native languages and cultural contexts. The dataset covers four core reasoning categories: language-specific linguistic reasoning, word games and riddles, cultural/traditional reasoning, and culture-related mathematical reasoning. The dataset construction process involves recruiting native speakers to create challenging reasoning questions, paired with objective and concise final answers to facilitate automatic evaluation. MultiNRC aims to address the current gaps in the multilingual reasoning capabilities of large language models (LLMs), and promote research on multilingual and culturally grounded evaluation.
MultiNRC: Multilingual Native Reasoning Challenge 数据集概述
数据集简介
MultiNRC是一个用于评估大型语言模型多语言推理能力的挑战性基准数据集,专注于法语、西班牙语和中文。数据集包含超过1,000个由母语者编写的推理问题,旨在捕捉语言和文化上的细微差别。
关键特性
- 支持语言:法语、西班牙语、中文
- 问题类别:
- 语言特定的语言推理
- 文字游戏与谜语
- 文化推理与传统
- 具有文化相关性的数学推理
- 英文等效内容:针对文化/传统和数学推理类别,提供人工翻译的英文版本以便直接比较
- 真实答案:每个提示都附带简短、客观的答案用于自动评估
数据结构
每个数据条目包含:
- 母语提示和答案(
i18n_prompt,i18n_gtfa) - (数学推理和文化推理类别任务)英文等效提示和答案(
english_prompt,english_gtfa) - 元数据:
task_id,language,category
数据集配置
- 默认配置:
- 测试集路径:
test/data-00000-of-00001.arrow
- 测试集路径:
- 数据规模:1K<n<10K
引用信息
bibtex @article{fabbri2025multinrc, title = {MultiNRC: A Challenging Native Multilingual Reasoning Evaluation Benchmark for LLMs}, author = {Fabbri, Alexander R. and Mares, Diego and Flores, Jorge and Mankikar, Meher and Hernandez, Ernesto and Lee, Dean and Liu, Bing and Xing, Chen}, year = {2025}, note = {arXiv preprint, arXiv:XXXX.XXXXX} }




