FTII-Bench
收藏资源简介:
FTII-Bench是由上海交通大学和上海高级算法研究所共同创建的一个综合多模态基准数据集,旨在评估大型视觉语言模型在图文插入任务中的表现。该数据集包含625篇高质量的中英文新闻文章,涵盖10个不同的新闻领域。数据集的创建过程包括从新华网和BBC新闻中手动收集数据,并设计了两种类型的问题:单选题和流插入题,以全面评估模型的多维度能力。FTII-Bench的应用领域主要集中在复杂的多模态任务评估,旨在解决现有基准在评估模型综合能力方面的不足。
FTII-Bench is a comprehensive multimodal benchmark dataset jointly developed by Shanghai Jiao Tong University and Shanghai Institute of Advanced Algorithms, which aims to evaluate the performance of large vision-language models on image-text insertion tasks. This dataset contains 625 high-quality Chinese and English news articles covering 10 distinct news categories. The development process of the dataset includes manually curating data from Xinhua News Agency and BBC News, as well as designing two types of questions: multiple-choice questions and stream insertion questions, to comprehensively evaluate the multi-dimensional capabilities of models. The primary application scope of FTII-Bench focuses on complex multimodal task evaluation, aiming to address the gaps in existing benchmarks when evaluating the comprehensive capabilities of models.
FTIIBench
数据集概述
- 名称: FTIIBench
- 描述: 这是一个名为 "FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion" 的官方代码仓库。
数据和代码
- 状态: 代码和数据将可用。

- 1FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion上海交通大学,上海高级算法研究所 · 2024年



