OpenEvents V1
收藏资源简介:
OpenEvents V1是一个大规模的基准数据集,旨在推动以事件为中心的视觉语言理解的发展。不同于传统的图像描述和检索数据集,OpenEvents V1专注于通过两个主要任务进行上下文和时态定位:(1)生成丰富的事件感知图像描述;(2)基于叙事风格文本查询检索事件相关的图像。该数据集包含来自CNN和The Guardian的超过20万篇新闻文章和40万张相关图像,涵盖了多个领域和时间跨度。我们为这两个任务提供了广泛的基线结果和标准化的评估协议。OpenEvents V1为开发能够对复杂现实世界事件进行深度推理的多模态模型奠定了坚实的基础。
OpenEvents V1 is a large-scale benchmark dataset designed to propel the advancement of event-centric visual language understanding. Unlike traditional image description and retrieval datasets, OpenEvents V1 focuses on contextual and temporal localization through two primary tasks: (1) generating rich event-aware image descriptions; and (2) retrieving event-related images based on narrative style text queries. The dataset encompasses over 200,000 news articles and 400,000 related images from CNN and The Guardian, spanning multiple domains and time periods. We provide a wide range of baseline results and standardized evaluation protocols for these two tasks. OpenEvents V1 lays a solid foundation for the development of multimodal models capable of performing deep inferences on complex real-world events.




