遇见数据集

Multimodal Dataset for Investment-Related Deceptive Content Detection on Social Platforms

收藏
Mendeley Data2026-05-21 收录
官方服务:

资源简介:

This dataset is a processed and anonymised dataset for research on deceptive or suspicious investment-related content in social platforms. The dataset integrates five public source groups covering phishing text, spam email, fake-profile posts, Twitter bot-detection records, and finance-related social media data, which were harmonised into a common binary schema and filtered for investment relevance. The deposited file contains 16,202 records and 32 columns. It includes source identifiers, text content, binary labels, deterministic data partitions, investment-filter outputs, and 20 standardized behavioral metadata features. These metadata features span raw account and content counts, relational proxy measures, and boolean profile indicators, enabling both text-only and multimodal analysis. The dataset contains 14,035 text-plus-metadata records and 2,167 text-only records, allowing evaluation under partial-modality conditions where behavioral information may be unavailable. The variables support studies in multimodal classification, metadata ablation, cross-source benchmarking, and analysis of behavioral patterns associated with deceptive investment-related communication. Labels represent investment-related deceptive or suspicious behaviour derived from harmonized source annotations after filtering, and should not be interpreted as verified ground-truth investment scam status for every individual record. This dataset is suitable for benchmarking machine learning models, studying heterogeneous social-platform deception signals, and supporting reproducible experiments on investment-focused content detection.

创建时间:
2026-05-13
二维码
社区交流群
二维码
科研交流群
商业服务