CrediBench
收藏资源简介:
CrediBench是一个大规模数据处理流程,用于构建时间网络图,以检测网络信息真伪。该数据集包含45百万个节点和1亿个边,是目前为止最大的网络图数据集,用于信息真伪研究。数据集通过Common Crawl存档中提取,包含文本内容和超链接结构。数据集适用于自然语言处理和图机器学习方法,以评估网络域的可信度。
CrediBench is a large-scale data processing pipeline for constructing temporal networks to verify the authenticity of online information. This dataset contains 45 million nodes and 100 million edges, making it the largest network dataset to date for research on information authenticity. Extracted from Common Crawl archives, it includes both textual content and hyperlink structures. This dataset is suitable for natural language processing and graph machine learning methods to evaluate the credibility of network domains.
CrediBench: Building Web-Scale Network Datasets for Information Integrity
数据集概述
CrediBench是一个用于构建时序网络数据集的大规模数据处理流程,专门用于信息完整性研究。该数据集通过联合建模文本内容和超链接结构来支持虚假信息检测。
核心特征
- 数据规模:包含4500万个节点和10亿条边
- 数据来源:基于2024年12月Common Crawl档案的一个月快照
- 数据类型:时序网络图,捕捉内容和网站间引用关系的动态演变
- 应用领域:虚假信息检测、信息完整性研究
技术特点
- 同时建模文本内容和超链接结构
- 捕捉一般虚假信息领域的动态演变
- 支持学习衡量来源可靠性的可信度分数
- 是目前公开可用的最大虚假信息研究网络图数据集
资源获取
- 数据处理管道和实验代码可通过提供的链接获取
- 数据集存储在指定文件夹中
相关领域
- 社会与信息网络 (cs.SI)
- 分布式、并行与集群计算 (cs.DC)
- 机器学习 (cs.LG)

- 1CrediBench: Building Web-Scale Network Datasets for Information IntegrityMila -Quebec AI Institute, McGill University, Concordia University, UC Berkeley, Université de Montréal, University of Oxford · 2025年



