CompreCap
收藏资源简介:
CompreCap基准以人类注释的场景图为特点,专注于综合图像字幕的评估。它提供了新的语义分割注释,平均掩码覆盖率为95.83%。此外,数据集还包括高质量的对象属性描述和对象之间的方向关系描述,构成一个完整的有向场景图结构。注释包括分割掩码、类别名称、属性描述和关系描述,存储在./anno.json文件中。基于此基准,研究人员可以全面评估大型视觉语言模型生成的图像字幕的质量。
The CompreCap benchmark is characterized by human-annotated scene graphs, focusing on comprehensive evaluation of image captioning. It provides novel semantic segmentation annotations with an average mask coverage of 95.83%. Additionally, the dataset includes high-quality object attribute descriptions and directional relationship descriptions between objects, forming a complete directed scene graph structure. The annotations, including segmentation masks, category names, attribute descriptions and relationship descriptions, are stored in the ./anno.json file. With this benchmark, researchers can comprehensively evaluate the quality of image captions generated by large vision-language models.
数据集卡片:CompreCap
数据集描述
CompreCap 基准数据集以人工标注的场景图为核心,专注于综合图像描述的评估。该数据集为图像中的常见对象提供了新的语义分割标注,平均掩码覆盖率为 95.83%。除了对对象的仔细标注外,CompreCap 还包括高质量的对象属性描述以及对象之间的方向性关系描述,构成了一个完整且有向的场景图结构。
分割掩码、类别名称、属性描述和关系描述的标注保存在 ./anno.json 文件中。基于 CompreCap 基准,研究人员可以全面评估大型视觉-语言模型生成的图像描述质量。评估代码可在 此处 获取。
许可信息
图像数据以标准的 Creative Common CC-BY-4.0 许可证分发,单个图像受其各自版权保护。
引用
BibTeX: bibtex @article{CompreCap, title={Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning}, author={Fan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu, Wei Zhai, Yang Cao, Yujun Shen, Zheng-Jun Zha}, booktitle={arXiv}, year={2024} }




