Text-ADBench
收藏资源简介:
Text-ADBench是一个基于大型语言模型(LLMs)嵌入的文本异常检测基准,涵盖了新闻、社交媒体和科学出版物等多个领域的文本数据集。该数据集包括早期语言模型(如GloVe、BERT)、多个LLMs(如LLaMa-2、LLama-3、Mistral、OpenAI的小型、ada和大型模型)的嵌入,并采用了三种池化策略(均值、序列结束标记和加权均值)。数据集包含八个真实世界的文本数据集,用于评估文本异常检测方法的性能。
Text-ADBench is a text anomaly detection benchmark based on embeddings of large language models (LLMs). It covers text datasets from multiple domains including news, social media, and scientific publications. This dataset includes embeddings from early language models such as GloVe and BERT, as well as multiple LLMs like LLaMa-2, LLaMa-3, Mistral, and OpenAI's small, ada, and large models. Three pooling strategies are adopted: mean pooling, sequence end token pooling, and weighted mean pooling. The benchmark contains eight real-world text datasets for evaluating the performance of text anomaly detection methods.
Text-ADBench 数据集概述
数据集简介
- 名称:Text-ADBench (Text Anomaly Detection Benchmark based on LLMs Embedding)
- 任务:文本异常检测
- 应用领域:欺诈检测、错误信息识别、垃圾邮件检测、内容审核等
- 论文链接:Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
数据集特点
-
嵌入模型多样性:
- 早期语言模型(GloVe, BERT)
- 多种大型语言模型(LLaMa-2, LLama-3, Mistral, OpenAI (small, ada, large))
-
多领域文本数据:
- 新闻、社交媒体、科学出版物等
-
评估指标:
- AUROC、AUPRC
数据集来源
- 20 Newsgroups: http://qwone.com/~jason/20Newsgroups/
- Reuters: https://raw.githubusercontent.com/nltk/nltk_data/gh-pages/packages/corpora/reuters.zi
- IMDB: http://ai.stanford.edu/~amaas/data/sentiment/
- SST-2: https://huggingface.co/datasets/stanfordnlp/sst2
- SMS Spam: https://huggingface.co/datasets/ucirvine/sms_spam
- Enron Emails: https://huggingface.co/datasets/Hellisotherpeople/enron_emails_parsed
- Web of Science: https://huggingface.co/datasets/river-martin/web-of-science-with-label-texts
- DBpedia: https://huggingface.co/datasets/fancyzhx/dbpedia_14
使用说明
-
环境要求:
- Python 3.8
- 安装依赖:
pip install requirements.txt
-
数据下载:
- 文本数据和文本嵌入可从 Text-ADBench 下载
-
配置:
- 修改
configs.py文件,设置有效的DATA_DIR和EMBEDDING_DIR
- 修改
-
功能模块:
- 文本嵌入:
./embedding/ - 异常检测:
./anomaly_detection/ - 低秩预测:
./low_rank_prediction/
- 文本嵌入:
引用
-
本工作: bibtex @misc{xiao2025textadbenchtextanomalydetection, title={Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding}, author={Feng Xiao and Jicong Fan}, year={2025}, eprint={2507.12295}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.12295}, }
-
嵌入模型:
- Llama-2-7B-chat
- Mistral-7B-Instruct-v0.2
- Llama-3-8B-Instruct
- LLM2Vec

- 1Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding香港中文大学(深圳) · 2025年



