Seeing Culture Benchmark (SCB)
收藏资源简介:
Seeing Culture Benchmark (SCB) 是一个旨在评估视觉语言模型 (VLM) 在文化推理方面的基准数据集。该数据集由来自七个东南亚国家的 1,065 张图像组成,这些图像涵盖了 138 种文化元素,分为五个类别:音乐、游戏、舞蹈、庆祝和婚礼。数据集还包含 3,178 个问题,其中 1,093 个是独特且经过人类标注者精心策划的。SCB 的创建过程涉及使用大型语言模型 ChatGPT 生成文化概念,并通过调查收集来自当地个体的反馈以验证这些概念。该数据集旨在解决 VLM 在处理跨文化场景时视觉推理和空间定位之间的差异问题。
Seeing Culture Benchmark (SCB) is a benchmark dataset designed to evaluate Vision-Language Models (VLMs) for cultural reasoning. This dataset comprises 1,065 images sourced from seven Southeast Asian countries, covering 138 distinct cultural elements and categorized into five categories: music, games, dance, celebrations, and weddings. The dataset also includes 3,178 questions, of which 1,093 are unique and carefully curated by human annotators. The development of SCB involved generating cultural concepts using ChatGPT, a Large Language Model (LLM), and collecting feedback from local participants via surveys to validate these concepts. This dataset aims to address the discrepancy between the visual reasoning and spatial localization capabilities of VLMs when handling cross-cultural scenarios.

- 1Seeing Culture: A Benchmark for Visual Reasoning and Grounding新加坡管理大学,万隆理工学院 · 2025年



