CopyGuard Benchmark Dataset
收藏资源简介:
该数据集由浙江大学等机构联合构建,旨在评估大型视觉语言模型(LVLM)对版权内容的识别与合规能力。数据集包含5万条多模态查询-内容对,覆盖书籍摘录、新闻文章、音乐歌词和代码文档四类版权材料,并模拟含版权声明与无声明两种现实场景。数据来源包括Goodreads畅销书、BBC/CNN新闻、Spotify热门音乐及Hugging Face/PyPI技术文档,通过严格的时间线和主题筛选确保多样性。其构建流程包含版权材料采集、声明添加和查询生成三个步骤,专门用于检测模型在重复、提取、改写和翻译四种侵权场景下的行为。该基准的建立为开发版权感知的多模态系统提供了重要支撑,助力解决AI生成内容引发的知识产权风险问题。
This dataset was jointly constructed by institutions including Zhejiang University, aiming to evaluate the copyright content recognition and compliance capabilities of Large Vision-Language Models (LVLM). The dataset contains 50,000 multimodal query-content pairs, covering four types of copyrighted materials: book excerpts, news articles, music lyrics, and code documentation, and simulates two realistic scenarios with and without copyright statements. Its data sources include bestsellers from Goodreads, BBC/CNN news, popular music from Spotify, and technical documentation from Hugging Face/PyPI, with diversity ensured through strict timeline and topic screening. The construction workflow consists of three steps: copyrighted material collection, statement addition, and query generation. It is specifically designed to detect model behaviors in four infringement scenarios: reproduction, extraction, rewriting, and translation. The establishment of this benchmark provides critical support for the development of copyright-aware multimodal systems, and helps address intellectual property risks arising from AI-generated content.
CopyGuard数据集概述
数据集基本信息
- 数据集名称:CopyGuard
- 托管平台:GitHub
- 访问地址:https://github.com/bluedream02/CopyGuard
数据集描述
根据README文件内容,该数据集未提供详细的功能、用途、数据内容、规模、格式、领域、应用场景、创建者、许可证、引用方式、更新历史、依赖项、使用方法或相关论文等具体信息。




