WebUIBench
收藏资源简介:
WebUIBench是一个大规模的综合基准,用于评估多模态大型语言模型在WebUI-to-Code任务上的性能。该数据集包括来自超过700个真实世界网站的超过21000个问题-答案对,涉及9个子任务。
WebUIBench is a large-scale comprehensive benchmark designed to evaluate the performance of multimodal large language models on the WebUI-to-Code task. This dataset includes over 21,000 question-answer pairs sourced from more than 700 real-world websites, covering 9 subtasks.
WebUIBench 数据集概述
基本信息
- 许可证: CC-BY-4.0
- 主页: https://github.com/MAIL-Tele-AI/WebUIBench
- 论文: WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
数据集简介
WebUIBench 是一个用于评估多模态大语言模型(MLLMs)在 WebUI-to-Code 任务中性能的大规模综合基准测试。数据集包含:
- 21K+ 问答对
- 0.7K+ 真实网站
- 9 个子任务
子任务配置
-
Element_Classification
- 特征: id, question, image_id, image, answer, subtask
- 测试集: 950 个样本
- 下载大小: 442,962,174 字节
-
Attribute_Regconition
- 特征: id, question, image_id, image, answer, subtask
- 测试集: 3,718 个样本
- 下载大小: 1,679,258,113 字节
-
Visual_Grounding
- 特征: id, question, image_id, image, answer, subtask
- 测试集: 3,934 个样本
- 下载大小: 1,897,962,456 字节
-
OCR
- 特征: id, question, image_id, image, answer, target_[x1,y1,x2,y2], subtask
- 测试集: 2,460 个样本
- 下载大小: 1,147,237,990 字节
-
Code_Error_Correction
- 特征: id, question, code_with_error, answer, subtask
- 测试集: 2,635 个样本
- 下载大小: 2,885,440 字节
-
Code_Function_Editing
- 特征: id, question, function_description, answer, subtask
- 测试集: 2,290 个样本
- 下载大小: 2,712,168 字节
-
Webpage_HTML_Matching
- 特征: id, question, image_id, image, answer, subtask
- 测试集: 2,143 个样本
- 下载大小: 1,003,289,265 字节
-
Webpage_HTML_Retrieval
- 特征: id, question, image_id, image, answer, subtask
- 测试集: 2,345 个样本
- 下载大小: 1,109,887,493 字节
联系方式
- Zhiyu Lin: zyllin@bjtu.edu.cn
- Zhengda Zhou: zhengdazhou@smail.nju.edu.cn
- Zhiyuan Zhao: tuzixini@gmail.com
引用
bibtex @article{xx, title={WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code}, author={xx}, journal={arXiv preprint arXiv:xx}, year={2025} }




