Overtone Global News Publications Journalistic Integrity Dataset
收藏资源简介:
Data uniqueness: we use custom built and trained NLP algorithms to assess qualitative metrics inherent in text content. We focus on what's in the text, not metadata such as publication or engagement. Our AI algorithms are co-created by NLP & journalism experts. Our datasets have all been human-reviewed and labeled. Dataset: CSV containing URL and/or body text, with attributed scoring as an integer and model confidence as a percentage. We ignore metadata such as author, publication, date, word count, shares and so on, to provide a clean and maximally unbiased assessment of content. Our data is provided in CSV/RSS/JSON format. One row = one scored article. CSV contains URL and/or body text, with attributed scoring as an integer and model confidence as a percentage. Integrity indicators provided as integers on a 1–5 scale. We also have custom models with 35 categories that can be added on request. Data sourcing: public websites, crawlers, scrapers, other partnerships where available. We generally can assess content behind paywalls as well as without paywalls. We source from ~4,000 news outlets, examples include: Bloomberg, CNN, BCC are one each. Countries: all English-speaking markets world-wide. Includes English-language content from non English majority regions, such as Germany, Scandinavia, Japan. Also available in Spanish on request. Use-cases: assessing the implicit integrity and reliability of an article. There is correlation between integrity and human value: we have shown that articles scoring highly according to our scales show increased, sustained, ongoing end-user engagement. Clients also use this to assess journalistic output, publication relevance and to create datasets of 'quality' journalism. Overtone provides a range of qualitative metrics for journalistic, newsworthy and long-form content. We find, highlight and synthesise content that shows added human effort and, by extension, added human value.
数据唯一性说明:我们采用自研并训练的自然语言处理(Natural Language Processing,NLP)算法,评估文本内容固有的质性指标。我们聚焦文本本身的内容,而非元数据(如发布信息、互动数据等)。本团队的AI算法由NLP与新闻学专家联合研发,所有数据集均经过人工审核与标注。 数据集格式:支持逗号分隔值(Comma-Separated Values,CSV)、简易信息聚合(Really Simple Syndication,RSS)、JavaScript对象表示法(JavaScript Object Notation,JSON)格式。其中CSV文件包含文章URL和/或正文文本,附带整数形式的归因评分,以及百分比形式的模型置信度。我们会忽略作者、发布平台、发布日期、字数、分享量等元数据,以实现对内容的纯净且尽可能无偏的评估。单条数据行对应一篇已评分文章。此外,我们还提供1-5级整数区间的完整性指标。若有需求,可启用包含35个类别的自定义模型。 数据来源:公开网站、爬虫抓取、现有合作渠道等。我们通常可评估付费墙内外的文章内容。我们的数据源覆盖约4000家新闻媒体,例如彭博社(Bloomberg)、美国有线电视新闻网(CNN)、BCC等。 覆盖市场:覆盖全球所有英语使用市场,包括非英语主流地区的英语内容,如德国、斯堪的纳维亚地区、日本等。此外可根据需求提供西班牙语版本数据。 应用场景:用于评估文章的隐性完整性与可信度。我们的研究表明,符合本平台评分标准的高评分文章,其终端用户参与度更高且更具持续性。客户还可利用该数据集评估新闻产出质量、媒体相关性,并构建「优质」新闻内容数据集。 Overtone可为新闻、时事及长文内容提供一系列质性评估指标。我们能够识别、筛选并整合那些投入了额外人力、进而具备更高人类价值的内容。




