官方服务:
资源简介:
Output of topic modeling in csv and html form. 400 iterations.
应用场景:
创建时间:
2012-04-23
相关数据集
google/civil_comments
该数据集包含来自Civil Comments平台的公开评论,该平台是一个独立新闻网站的评论插件。这些评论创建于2015年至2017年之间,出现在全球大约50个英文新闻网站上。当Civil Comments在2017年关闭时,他们选择将这些公开评论保存在一个持久的开放档案中,以便未来的研究使用。原始数据包括公开评论文本、一些相关的元数据(如文章ID、时间戳和评论者生成的“文明”标签),但不包括用户I
Hugging Face2024-01-25 更新1250
sdananya/wiki_data_with_label_chunk_71
该数据集包含了与文章相关的信息,如文章标题(title)、正文内容(text)、URL链接(url)、维基百科ID(wiki_id)、浏览量(views)、段落ID(paragraph_id)、语言类型(langs)、嵌入表示(emb)、关键词(keywords)、标签(labels)和分类(categories)。数据集主要被用于训练,包含1000个示例。不过具体的应用场景和详细内容描述并未在R
Hugging Face2025-02-13 更新90
DISEASES v2 (human, text mining, filtered)
This file contains the filtered non-redundant set of gene–disease associations obtained from automatic text mining in DISEASES v2.
DataCite Commons2025-06-01 更新80
DECM Annotated Corpus
The DECM Corpus is a digital corpus of the texts of Relaciones Geográficas de Nueva España (the Geographic Reports of New Spain) with different versions, including a machine ready version, a gold stan
DataCite Commons2022-12-08 更新120
Marcus2112/minipile_density-proportioned_tiny
这是一个基于The Pile Deduplicated的数据集,包含文本(text)和索引(pile_idx)两种特征。数据集被划分为训练集(train)、验证集(validation)和测试集(test),分别包含842,967、500和10,000个示例。整个数据集大小为5,279,901,670字节,下载大小为2,903,588,242字节。数据集使用英语。
Hugging Face2025-01-23 更新80



