这是一个不断增长的AI文本检测预测数据集,包含来自AI Text Detector Space平台的预测结果和用户反馈。每次在平台上分析文本或URL时,预测结果都会被记录到数据集中。用户还可以提供反馈(正确或错误),这些反馈与预测结果一起存储。数据集格式为JSONL,每个记录包含唯一ID、分析的文本、URL(如果直接粘贴则为空)、模型的预测(ai或human)、置信度分数、用户反馈(如果有)以及时
Real vs AI Corpus是一个大规模二进制图像分类数据集,专为训练AI图像检测器而设计。该数据集由17个公开的HuggingFace资源构建而成,所有资源均以流式合并方式处理,无需中间本地存储。数据集包含真实图像和AI生成图像,用于区分真实图像和AI生成图像。数据集的使用、架构、来源及许可证信息均在README中有详细说明。
We built an AI text recognition dataset that contains both human text and ChatGPT generated text in a total quantity of 35K. It contains three categories, namely news text, comment text, and question
The advent of ChatGPT in November 2022 immediately rang alarm bells about the integrity of academic assessment, especially in distance learning. The challenge of detecting generative AI (GenAI) misuse