ARB
收藏资源简介:
ARB数据集是一个全面的阿拉伯语多模态推理基准,包含11个不同领域的1,356个样本和5,119个推理步骤,涵盖文本和视觉模态,强调逐步推理过程,由母语阿拉伯语人士和领域专家审核,支持阿拉伯推理和多模态AI研究。
The ARB dataset is a comprehensive Arabic multimodal reasoning benchmark that includes 1,356 samples across 11 distinct domains and 5,119 reasoning steps, covering both textual and visual modalities with an emphasis on step-by-step reasoning processes. The dataset has been audited by native Arabic speakers and domain experts, and it supports Arabic reasoning and multimodal AI research.
ARB: 综合性阿拉伯多模态推理基准数据集
数据集概述
- 名称: ARB (A Comprehensive Arabic Multimodal Reasoning Benchmark)
- 类型: 多模态推理基准数据集
- 语言: 阿拉伯语
- 样本数量: 1,356个多模态样本
- 推理步骤数量: 5,119个精心策划的推理步骤
- 领域覆盖: 11个多样化领域
关键特性
- 强调逐步推理,超越最终答案预测
- 每个样本包含2-6+个推理步骤链,与人类逻辑一致
- 由阿拉伯语母语者和领域专家策划和验证
- 来源包括原始阿拉伯数据、高质量翻译和合成样本
- 提供强大的评估框架,衡量最终答案准确性和推理质量
数据集结构
特征
image: 图像输入question: 阿拉伯语推理提示answer: 最终解决方案(阿拉伯语)choices: MCQ选项steps: 有序推理链domain: 领域类别Curriculum: 课程类别
分割
preview: 20个示例,6,553,087字节train: 1,355个示例,657,252,987.185字节
评估协议
- 评估方法:
- 词法和语义相似性评分:BLEU、ROUGE、BERTScore、LaBSE
- 使用LLM-as-Judge的逐步评估
- 评估因素: 包括忠实度、解释深度、连贯性、幻觉等10个因素
评估结果
闭源模型
| 指标/模型 | GPT-4o | GPT-4o-mini | GPT-4.1 | o4-mini | Gemini 1.5 Pro | Gemini 2.0 Flash |
|---|---|---|---|---|---|---|
| 最终答案 (%) | 60.22 | 52.22 | 59.43 | 58.93 | 56.70 | 57.80 |
| 推理步骤 (%) | 64.29 | 61.02 | 80.41 | 80.75 | 64.34 | 64.09 |
开源模型
| 指标/模型 | Qwen2.5-VL | LLaMA-3.2 | AIN | LLaMA-4 Scout | Aya-Vision | InternVL3 |
|---|---|---|---|---|---|---|
| 最终答案 (%) | 37.02 | 25.58 | 27.35 | 48.52 | 28.81 | 31.04 |
| 推理步骤 (%) | 64.03 | 53.20 | 52.77 | 77.70 | 63.64 | 54.50 |
引用
bibtex @misc{ghaboura2025arbcomprehensivearabicmultimodal, title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark}, author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer}, year={2025}, eprint={2505.17021}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.17021}, }




