virgool_62k
收藏资源简介:
该数据集是从virgool.io网站上公开收集的数据集合,通过特定标签和用户策略性地提取。数据集包含约62,000条记录,涵盖标题、文本、标签、点赞数、回复数、阅读时间、用户ID和URL等关键属性。此资源特别有利于研究人员和开发者预训练大型语言模型(LLMs),因为'文本'列提供了丰富的语言使用语料库。此外,'标签'列非常适合主题建模应用。'点赞'和'回复'列提供了量化见解,可用于衡量内容参与度,并有助于开发识别高质量或信息丰富内容的分类器。
This dataset is a publicly accessible collection of data strategically extracted from the website virgool.io using specific tags and user-oriented selection strategies. It contains approximately 62,000 records, covering key attributes including title, text content, tags, like count, reply count, reading time, user ID, and URL. This resource is particularly beneficial for researchers and developers to pre-train Large Language Models (LLMs), as the 'text' column provides a rich corpus of linguistic usage. Furthermore, the 'tags' column is highly suitable for topic modeling applications. The 'like' and 'reply' columns offer quantitative insights that can be used to measure content engagement, and aid in developing classifiers for identifying high-quality or informative content.
数据集概述
数据集描述
- 数据来源: 从 virgool.io 网站爬取的公开可用数据。
- 数据量: 约 62,000 条记录。
- 字段信息: 包含标题、文本、标签、点赞数、回复数、阅读时间、用户ID和URL。
- 语言: 波斯语。
- 许可证: Apache-2.0。
数据集结构
数据集包含以下8个字段:
- title: 标题
- text: 文本
- tags: 标签
- likes: 点赞数
- replies: 回复数
- reading_time: 阅读时间
- user_id: 用户ID
- URL: 链接
偏见、风险和限制
该数据集包含来自 virgool.io 上不同博主的个人观点,信息可能不总是事实性的,并可能反映作者的个人偏见。




