VisualWebInstruct
收藏资源简介:
VisualWebInstruct是一个通过搜索引擎创建的多学科、高质量数据集,旨在解决推理型多模态数据集的稀缺问题。该数据集包含约90万对问答对,其中40%为视觉问答对,其余为文本问答对。通过内容提取、过滤和合成的管道,从超过70万个独特的URL源中收集和处理HTML数据。在多个基准测试中,使用VisualWebInstruct微调的模型表现出显著的性能提升。
VisualWebInstruct is a multidisciplinary, high-quality dataset developed via search engines, aimed at addressing the scarcity of reasoning-oriented multimodal datasets. It contains approximately 900,000 question-answer pairs, 40% of which are visual question-answer pairs, with the remainder being text-based question-answer pairs. HTML data was collected and processed from over 700,000 unique URL sources through a pipeline of content extraction, filtering and synthesis. Models fine-tuned with VisualWebInstruct have demonstrated significant performance improvements across multiple benchmark tests.
VisualWebInstruct 数据集概述
数据集简介
- 数据集名称:VisualWebInstruct
- 创建目的:为提高视觉语言模型在推理任务上的性能,解决高质量、多样化训练数据的稀缺问题。
- 数据来源:通过搜索引擎(Google Image)搜索与精选的30,000个种子图像相似的网站,收集并处理超过700K个唯一URL来源的HTML内容。
- 数据构成:包含约900K个问题-答案对,其中40%为视觉问答对,其余为文本问答对。
数据集特点
- 学科覆盖:涵盖数学、物理、金融、化学等多个学科。
- 性能提升:在VisualWebInstruct上微调的模型显示出显著的性能提升,例如Llava-OV-mid训练的模型在各项基准测试中提高10-20%,MAmmoTH-VL训练的模型提高5%。
- 最佳模型表现:MAmmoTH-VL2模型在10B参数类别中,在MMMU-Pro-std(40.7%)、MathVerse(42.6%)和DynaMath(55.7%)上展示出最先进性能。
引用信息
@article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} }




