遇见数据集

DETECT-AI-Dataset

收藏
Hugging Face2026-03-14 更新2026-03-16 收录
官方服务:

资源简介:

DETECT-AI 多模态 AI 内容检测数据集是一个用于检测 AI 生成内容的大规模数据集,涵盖文本、图像、视频和音频四种模态,支持包括英语、中文、阿拉伯语等在内的 19 种语言。数据集每月从 19 个全球来源爬取超过 10 亿条经过验证的样本,并通过 8 个专业 AI 检测模型的加权集成进行标注。数据标注分为三类:AI_GENERATED(AI 生成,置信度 ≥ 0.75)、HUMAN(人类创作,置信度 ≤ 0.35)和 UNCERTAIN(不确定,置信度 0.35–0.75)。数据集采用 CC-BY-4.0 许可,允许研究和商业用途。数据组织结构清晰,包含元数据、样本内容和处理日志等,适用于文本分类、图像分类、视频分类和音频分类等多种任务。

The DETECT-AI Multimodal AI-Generated Content Detection Dataset is a large-scale dataset designed for detecting AI-generated content, covering four modalities including text, image, video and audio, and supporting 19 languages such as English, Chinese and Arabic. It crawls over 1 billion verified samples monthly from 19 global sources, and is annotated via a weighted ensemble of 8 professional AI detection models. The dataset’s annotations are divided into three categories: AI_GENERATED (AI-generated, confidence score ≥ 0.75), HUMAN (human-created, confidence score ≤ 0.35) and UNCERTAIN (uncertain, confidence score ranging from 0.35 to 0.75). The dataset is licensed under CC-BY-4.0, permitting both research and commercial uses. With a well-organized data structure that includes metadata, sample content, processing logs and other relevant materials, it is suitable for various tasks such as text classification, image classification, video classification and audio classification.

创建时间:
2026-03-14
二维码
社区交流群
二维码
科研交流群
商业服务