CrediBench
收藏资源简介:
CrediBench是一个大规模数据处理流程,用于构建时间网络图,这些图共同模拟文本内容和超链接结构,用于检测网络上的不实信息。从2024年12月Common Crawl存档中提取的一个月快照包含4500万个节点和10亿条边,是迄今为止公开可用的最大的网络图数据集,用于不实信息研究。该数据集通过构建包含节点文本内容和结构丰富特征的时序图,为自然语言处理技术和图机器学习方法提供了应用机会,以评估网络域的可信度。
CrediBench is a large-scale data processing pipeline for constructing temporal network graphs, which collectively simulate text content and hyperlink structures to detect online misinformation. A one-month snapshot extracted from the December 2024 Common Crawl archive contains 45 million nodes and 1 billion edges, representing the largest publicly available network graph dataset for misinformation research to date. By constructing temporal graphs containing node text content and structurally rich features, this dataset enables the application of natural language processing (NLP) techniques and graph machine learning methods to evaluate the credibility of online domains.

- 1CrediBench: Building Web-Scale Network Datasets for Information IntegrityMila - Quebec AI Institute, McGill University, Concordia University, UC Berkeley, Université de Montréal, University of Oxford, AITHYRA · 2025年



