Hyperphantasia
收藏资源简介:
Hyperphantasia是一个合成基准,旨在评估多模态大型语言模型(MLLMs)的内心可视化能力。该数据集由四类精心设计的谜题组成,每个谜题都有三个难度级别,总共包含1200个样本。这些谜题旨在测试模型在推理、预测和抽象等任务中构建和操作内部视觉表示的能力。数据集已公开,可用于评估当前MLLMs在内心可视化能力方面的表现,并探索强化学习在提高视觉模拟能力方面的潜力。
Hyperphantasia is a synthetic benchmark developed to evaluate the mental visualization capabilities of multimodal large language models (MLLMs). This dataset comprises four categories of meticulously crafted puzzles, each with three difficulty levels, totaling 1200 samples. These puzzles are designed to test the model's capacity to construct and manipulate internal visual representations across tasks such as reasoning, prediction, and abstraction. The dataset is publicly available, allowing researchers to evaluate the performance of current MLLMs in terms of mental visualization capabilities and explore the potential of reinforcement learning in improving visual simulation abilities.
Hyperphantasia数据集概述
基本信息
- 许可证: MIT
- 数据集名称: Hyperphantasia
- 标签: mental_visualization, text, image, benchmark, puzzles
- 任务类别: multiple-choice, question-answering, visual-question-answering
- 语言: 英语 (en)
- 数据规模: 1K<n<10K
数据集描述
Hyperphantasia是一个合成的视觉问答(VQA)基准数据集,用于从视觉角度探究多模态大型语言模型(MLLMs)的心理可视化能力。该数据集揭示了最先进的模型在需要视觉模拟和想象的简单任务中表现不佳。
数据集内容
- 样本数量: 1200个
- 谜题类型: 4种不同的谜题
- 分类: 插值(Interpolation)和外推(Extrapolation)两类
- 难度级别: 三个难度级别,用于评估MLLMs心理可视化能力的范围和泛化性
使用信息
- 评估代码: 可在Github仓库中找到




