Bend the Truth
收藏资源简介:
该数据集包含六个不同领域(技术、教育、商业、体育、政治、娱乐)的新闻,其中真实新闻来自巴基斯坦、印度、英国和美国的主流新闻网站,而假新闻则是由专业记者编写的真实新闻的假版本。
This dataset comprises news articles from six distinct domains (technology, education, business, sports, politics, entertainment). The authentic news articles are sourced from mainstream news websites in Pakistan, India, the United Kingdom, and the United States, whereas the fake news articles are fabricated versions of the real news, crafted by professional journalists.
Urdu Fake News Dataset 概述
数据集介绍
- 名称: Urdu Fake News Dataset
- 内容: 包含5个不同领域的新闻数据,分别是体育、健康、技术、娱乐和商业。
- 真实新闻收集方法: 结合手动方法收集。
- 虚假新闻收集方法: 通过专业记者的众包注释收集。
数据集结构
- 数据集名称: "Bend the Truth"
- 领域: 技术、教育、商业、体育、政治、娱乐。
- 来源: 主要来自巴基斯坦、印度、英国和美国的多个主流新闻网站,如BBC Urdu News, CNN Urdu等。
- 结构: 包含两个文件夹,分别存放真实和虚假新闻,共5种新闻类型。
- 类别分布:
- 虚假新闻: 400条
- 真实新闻: 500条
引用信息
-
引用格式:
@article{MaazUrdufake2020, author = {Amjad, Maaz and Sidorov, Grigori and Zhila, Alisa and Gómez-Adorno, Helena and Voronkov, Ilia and Gelbukh, Alexander}, title = {Bend the Truth: A Benchmark Dataset for Fake News Detection in Urdu and Its Evaluation}, journal={Journal of Intelligent & Fuzzy Systems}, volume={39}, number={2}, pages={2457-2469}, doi = {10.3233/JIFS-179905}, year={2020}, publisher={IOS Press} }




