遇见数据集

ytaek-oh/eqben-images

收藏
Hugging Face2024-01-25 更新2024-03-04 收录
官方服务:

资源简介:

--- license: apache-2.0 --- <p align="center"> <h3 align="center"><a href="https://arxiv.org/abs/2303.14465" target='_blank'> <strong>Equivariant Similarity for Vision-Language Foundation Models</strong> </a></h3> <h2 align="center">ICCV 2023</h2> <p align="center"> <a href="https://scholar.google.com/citations?hl=en&user=wFduC9EAAAAJ" target='_blank'>Tan Wang</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=LKSy1kwAAAAJ" target='_blank'>Kevin Lin</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=WR875gYAAAAJ" target='_blank'>Linjie Li</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=legkbM0AAAAJ" target='_blank'>Chung-Ching Lin</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=rP02ve8AAAAJ" target='_blank'>Zhengyuan Yang</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=YG0DFyYAAAAJ" target='_blank'>Hanwang Zhang</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=bkALdvsAAAAJ" target='_blank'>Zicheng Liu</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=cDcWXuIAAAAJ" target='_blank'>Lijuan Wang</a> <br> Nanyang Technological University, &nbsp; Microsoft Corporation </p> </p> <br /><br /> # About This study explores the concept of equivariance in vision-language foundation models (VLMs), focusing specifically on the multimodal similarity function that is not only the major training objective but also the core delivery to support downstream tasks. Unlike the existing image-text similarity objective which only categorizes matched pairs as similar and unmatched pairs as dissimilar, equivariance also requires similarity to vary faithfully according to the semantic changes. Our key contributions are three-fold: 1. A novel benchmark named **EqBen** (Equivariant Benchmark) to benchmark VLMs with **visual-minimal change** samples. 2. A plug-and-play regularization loss **EqSim** (Equivariant Similarity Learning) to improve the equivariance of current VLMs. 3. Toolkit provides an **one-stop evaluation**: not only for EqBen, but also for previous related benchmarks (Winoground, VALSE, etc).<br> # Data Download - **Download images from huggingface hub:** Please check the Files and versions tab above. - Full-Test Set: the user can download the EqBen raw **[image data](https://drive.google.com/file/d/1e608uhd36ak_v7SnlMVaYcekBc4gBqzn/view?usp=drive_link)** (tar.gz file, ~100G) and [**annotation (after randomize)**](https://drive.google.com/file/d/1-CWEuZ5F0KQ4d94Y9rRtBsMIcqb8V7nm/view?usp=sharing) (200M) via Google Drive. **[UPDATE-2023-09]** The original annotation is the annotation after randomize (non-public) for the total fairness. And the users are required to upload the results json/np file to CodaLab for getting the final results. Due to the unstability of CodaLab, we decide to public the whole original annotation. This annotation file formalized similar to *Winoground* and can be downloaded [**here**](https://drive.google.com/file/d/1gNR4K2Cv4rbnjVRdHuBV6PuoJ5MlRXnZ/view?usp=sharing). - **Light** Full-Test Set: to improve the usability, we also provide a light version of EqBen by converting all the png image to the jpg using `convert`. Feel free to download [here](https://entuedu-my.sharepoint.com/:u:/g/personal/tan317_e_ntu_edu_sg/EcHBRcch6KREvzvGgrN67FMBUSVV4QPTQUiew0bxjcitFw?e=xiJiYL). But please note that you may make some small revisement to the path in the annotation (change the `.png` to `.jpg`). - Sub-Test Set: we also provide a 10% subset (~25K image-text pairs) for the ease of visualization and validation. The label of the EqBen sub-set is **opensource** and the **format follows the winoground style**. But please note that the samples in the subset is **randomly sorted** and not be classified to each category. Please down the raw **[image data](https://drive.google.com/file/d/13Iuirsvx34-9F_1Mjhs4Dqn59yokyUjy/view?usp=sharing)** (tar.gz file, ~10G) and [**annotation**](https://drive.google.com/file/d/18BSRf1SnBtGiEc42mzRLirXaBLzYE5Tt/view?usp=sharing) via Google Drive. --- * This is the unofficial distribution of images in the eqben benchmark. * For the official repository, please visit [https://github.com/Wangt-CN/EqBen](https://github.com/Wangt-CN/EqBen). * Some part of this README.md is taken from the official repository

license: apache-2.0 <p align="center"> <h3 align="center"><a href="https://arxiv.org/abs/2303.14465" target='_blank'> <strong>面向视觉语言基础模型的等变相似度学习</strong> </a></h3> <h2 align="center">ICCV 2023</h2> <p align="center"> <a href="https://scholar.google.com/citations?hl=en&user=wFduC9EAAAAJ" target='_blank'>王坦</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=LKSy1kwAAAAJ" target='_blank'>林凯文</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=WR875gYAAAAJ" target='_blank'>李林杰</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=legkbM0AAAAJ" target='_blank'>林仲靖</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=rP02ve8AAAAJ" target='_blank'>杨正远</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=YG0DFyYAAAAJ" target='_blank'>张汉旺</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=bkALdvsAAAAJ" target='_blank'>刘志成</a>,&nbsp; <a href="https://scholar.google.com/citations?hl=en&user=cDcWXuIAAAAJ" target='_blank'>王丽娟</a> <br> 南洋理工大学,&nbsp;微软公司 </p> </p> <br /><br /> # 研究概况 本研究探讨视觉语言基础模型(Vision-Language Foundation Models, VLMs)中的等变性概念,重点关注多模态相似度函数——该函数不仅是模型的核心训练目标,亦是支撑下游任务的关键输出模块。现有图文相似度目标仅将匹配样本判定为相似、非匹配样本判定为相异,而等变性要求相似度能够忠实响应语义变化。本研究的核心贡献包含三点: 1. 提出全新基准测试集**EqBen(等变基准测试集,Equivariant Benchmark)**,通过视觉最小变化样本对视觉语言基础模型进行基准测试。 2. 提出即插即用的正则化损失函数**EqSim(等变相似度学习,Equivariant Similarity Learning)**,用于提升现有视觉语言基础模型的等变性。 3. 提供一站式评测工具包:不仅可用于EqBen的评测,还支持此前相关基准测试集(如Winoground、VALSE等)的评测。<br> # 数据下载 - **从Hugging Face Hub下载图像**:请查看页面上方的“Files and versions”标签页。 - **完整测试集**:用户可通过Google Drive下载EqBen原始**[图像数据](https://drive.google.com/file/d/1e608uhd36ak_v7SnlMVaYcekBc4gBqzn/view?usp=drive_link)**(tar.gz格式,约100GB)与**[经随机化处理的标注文件](https://drive.google.com/file/d/1-CWEuZ5F0KQ4d94Y9rRtBsMIcqb8V7nm/view?usp=sharing)**(200MB)。 **[2023-09更新]** 原始标注为经随机化处理的标注(非公开),以保证整体公平性。此前用户需将评测结果的json/np文件上传至CodaLab以获取最终评测结果。鉴于CodaLab平台存在不稳定性,我们决定公开全部原始标注文件。该标注文件的格式与*Winoground*类似,可通过[此链接](https://drive.google.com/file/d/1gNR4K2Cv4rbnjVRdHuBV6PuoJ5MlRXnZ/view?usp=sharing)下载。 - **轻量版完整测试集**:为提升数据集易用性,我们通过`convert`工具将所有PNG图像转换为JPG格式,提供EqBen的轻量版本。可通过[此链接](https://entuedu-my.sharepoint.com/:u:/g/personal/tan317_e_ntu_edu_sg/EcHBRcch6KREvzvGgrN67FMBUSVV4QPTQUiew0bxjcitFw?e=xiJiYL)自由下载。但请注意:您需要对标注文件中的图像路径进行小幅修改,将`.png`后缀更改为`.jpg`。 - **子测试集**:我们还提供了10%的子集(约25K个图文对),方便可视化与验证。EqBen子集的标注**已开源**,格式遵循Winoground风格。但请注意:子集内的样本为随机排序,未按类别划分。请通过Google Drive下载原始**[图像数据](https://drive.google.com/file/d/13Iuirsvx34-9F_1Mjhs4Dqn59yokyUjy/view?usp=sharing)**(tar.gz格式,约10GB)与**[标注文件](https://drive.google.com/file/d/18BSRf1SnBtGiEc42mzRLirXaBLzYE5Tt/view?usp=sharing)**。 --- * 本分发包为EqBen基准测试集中图像的非官方分发版本。 * 官方仓库请访问[https://github.com/Wangt-CN/EqBen](https://github.com/Wangt-CN/EqBen)。 * 本README.md的部分内容源自官方仓库。

提供机构:
ytaek-oh
原始信息汇总

数据集概述

关于

本研究探索了视觉-语言基础模型(VLMs)中的等变性概念,特别关注于多模态相似性函数,这不仅是主要的训练目标,也是支持下游任务的核心交付内容。与现有的仅将匹配对分类为相似、不匹配对分类为不相似的图像-文本相似性目标不同,等变性还要求相似性根据语义变化忠实地变化。我们的主要贡献有三点:

  1. 一个名为 EqBen(等变性基准)的新基准,用于评估具有视觉最小变化样本的VLMs。
  2. 一个即插即用的正则化损失 EqSim(等变性相似性学习),以提高当前VLMs的等变性。
  3. 工具包提供了一个一站式评估:不仅适用于EqBen,还适用于先前的相关基准(如Winoground、VALSE等)。

数据下载

  • 完整测试集:用户可以通过Google Drive下载EqBen原始的**图像数据(tar.gz文件,约100G)和随机化后的标注**(200M)。
  • 轻量版完整测试集:为了提高可用性,我们还提供了一个轻量版的EqBen,通过将所有png图像转换为jpg格式。可以在此处下载:轻量版数据。请注意,您可能需要对标注中的路径进行一些小的修改(将.png改为.jpg)。
  • 子测试集:我们还提供了一个10%的子集(约25K图像-文本对),以便于可视化和验证。EqBen子集的标签是开源的,格式遵循Winoground风格。请注意,子集中的样本是随机排序的,并未分类到各个类别。请在此处下载原始的**图像数据(tar.gz文件,约10G)和标注**。

搜集汇总
数据集介绍
ytaek-oh/eqben-images 数据集图片
构建方式
在视觉-语言基础模型研究领域,EqBen数据集的构建聚焦于评估模型的等变性能力。该数据集通过精心设计视觉最小变化样本,系统性地收集了图像-文本对。构建过程涉及从多样化来源筛选原始图像,并基于语义细微差异生成对应的文本描述,确保每个样本对在视觉上仅存在微小改动,而文本则精确反映这种语义变化。数据标注遵循严谨的协议,以支持对模型相似度函数等变性的量化分析。
使用方法
使用EqBen数据集时,研究人员可从HuggingFace平台下载图像数据,并结合提供的标注文件进行模型评估。数据集支持一站式评估流程,用户可通过官方工具包计算模型在等变性任务上的性能指标。具体而言,需将模型预测的相似度结果整理为指定格式的JSON或NP文件,并提交至评估系统获取分数。对于快速验证,建议使用公开标注的子集进行初步测试,确保模型实现正确后再扩展到完整测试集。
背景与挑战
背景概述
随着视觉-语言基础模型在跨模态理解任务中展现出卓越性能,其内部相似性度量的鲁棒性与语义一致性成为研究焦点。由南洋理工大学与微软研究院团队于2023年构建的EqBen数据集,旨在系统评估模型对视觉细微变化的等变性响应能力。该数据集通过精心设计的视觉最小变化样本,挑战传统图像-文本匹配范式,推动相似性函数从静态匹配向动态语义感知演进,为ICCV 2023收录的等变相似性研究提供了基准支撑。
当前挑战
EqBen数据集致力于解决视觉-语言基础模型中相似性度量的语义敏感性问题,其核心挑战在于如何量化模型对视觉细微变化的等变响应能力。在构建过程中,需精确控制图像语义的渐进式变化,同时保持文本描述的对应性,这对样本对的语义对齐与变化梯度设计提出了极高要求。此外,大规模高质量数据标注的复杂性,以及评估指标与传统基准的兼容性整合,均是实现可靠等变性评估的关键难点。
常用场景
经典使用场景
在视觉-语言基础模型的研究领域,EqBen数据集以其视觉最小变化样本的独特设计,为评估模型对语义细微差异的敏感性提供了经典场景。该数据集通过精心构造的图像-文本对,要求模型不仅识别匹配与否,还需捕捉语义变化导致的相似度波动,从而深入检验模型在复杂多模态理解中的表现。这一场景常被用于基准测试,推动模型超越传统二元分类,迈向更精细的语义感知。
解决学术问题
EqBen数据集致力于解决视觉-语言基础模型中相似性函数的等变性缺失问题。传统方法仅将匹配对视为相似、非匹配对视为不相似,忽视了语义变化对相似度的连续影响。该数据集通过引入视觉最小变化样本,促使模型学习更忠实的相似度响应,从而提升模型在语义推理、细粒度对齐等任务上的性能,为多模态表示学习提供了新的理论框架与评估标准。
实际应用
在实际应用中,EqBen数据集可服务于多模态系统的鲁棒性优化与性能验证。例如,在图像检索、自动标注和视觉问答系统中,模型需对细微的视觉或文本变化保持敏感。该数据集帮助开发者测试模型在真实世界复杂场景下的稳定性,确保其输出符合人类语义直觉,进而提升智能助手、内容审核等应用的准确性与可靠性。
数据集最近研究
最新研究方向
在视觉-语言基础模型领域,ytaek-oh/eqben-images数据集作为EqBen基准的核心组成部分,正推动着对模型等变性能力的前沿探索。该数据集通过精心构建的视觉最小变化样本,旨在评估模型在语义细微变动下相似性度量的忠实性,从而揭示现有模型在理解复杂多模态关系时的局限性。当前研究热点集中于利用EqSim等正则化损失提升模型的等变性能,这一方向不仅与Winoground、VALSE等基准形成互补,更在图像生成、跨模态检索等下游任务中展现出深远影响,为构建更具鲁棒性和可解释性的多模态人工智能系统奠定了关键基础。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务