adaptive-classifier/ai-detector-data
收藏资源简介:
这是一个不断增长的AI文本检测预测数据集,包含来自AI Text Detector Space平台的预测结果和用户反馈。每次在平台上分析文本或URL时,预测结果都会被记录到数据集中。用户还可以提供反馈(正确或错误),这些反馈与预测结果一起存储。数据集格式为JSONL,每个记录包含唯一ID、分析的文本、URL(如果直接粘贴则为空)、模型的预测(ai或human)、置信度分数、用户反馈(如果有)以及时间戳。数据集持续更新,每次预测后都会添加新记录。
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space. Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click Correct or Incorrect to provide feedback, which gets stored alongside the prediction. The dataset is in JSONL format, with each record containing a unique ID, the analyzed text, URL (empty if pasted directly), the models prediction (ai or human), confidence score, user feedback (if any), and timestamp. The dataset updates live, with new records added on every Space inference.
数据集概述:ai-detector-data
基本信息
- 数据集名称: ai-detector-data
- 发布者: adaptive-classifier
- 许可证: Apache-2.0
- 任务类型: 文本分类
- 数据模态: 文本
- 数据格式: JSON
- 语言: 英语
- 数据集大小: 小于 1K 条记录(当前预览显示 584 行)
- 标签: ai-detection, ai-generated-text, human-vs-ai, text-classification, continuous-learning
- 相关库: Datasets, pandas, Polars
数据集结构
- 子集(Subset): default(584 行)
- 数据拆分(Split): train(584 行)
数据模式(Schema)
| 字段名 | 类型 | 说明 |
|---|---|---|
| id | string | 唯一标识符 |
| text | string | 文本内容 |
| url | string | 来源 URL(可选) |
| prediction | string | 预测结果("human"或"ai") |
| confidence | float64 | 预测置信度(范围 0.5 至 0.67) |
| feedback | string | 用户反馈("correct"、"incorrect"或 null) |
| timestamp | datetime | 记录生成时间戳 |
数据内容概述
该数据集是一个持续增长的 AI 文本检测预测结果集合,每条记录包含以下关键信息:
- 文本样本:涵盖新闻、社交媒体帖子、学术摘要、技术文章、虚构故事等多种类型的文本内容。
- 预测标签:每条记录标注为 "human"(人类撰写)或 "ai"(AI 生成)。
- 置信度分数:预测模型的置信度,范围通常在 0.5 至 0.67 之间。
- 用户反馈:部分记录包含用户提供的 "correct"(正确)或 "incorrect"(错误)反馈。
- 来源 URL:部分文本附带原始网页 URL。
数据用途与更新机制
- 数据来源:每条预测记录来自 AI Text Detector Space,每当用户在该空间分析文本或 URL 时,预测结果会自动追加到该数据集中。
- 反馈机制:用户可通过点击 "Correct" 或 "Incorrect" 提供反馈,反馈信息与预测记录一同存储。
- 持续更新:数据集是不断增长的集合,新记录会随时间持续添加。




