community-science-merged
收藏资源简介:
该数据集包含多个字段,如arxiv_id、reached_out_link、reached_out_success等,涉及到的数据类型包括字符串、浮点数和布尔值。数据集主要用于记录与学术论文相关的信息,包括论文的识别号、外部链接、成功联系的情况、笔记、模型数量、数据集数量、空间数量、标题、GitHub信息、GitHub星数、会议名称、点赞数、评论数、GitHub提及HF的情况、是否有制品、提交者和日期。数据集分为训练集,包含5064个样本,总大小为1127665字节。
This dataset contains multiple fields including arxiv_id, reached_out_link, reached_out_success, and others, with data types spanning strings, floating-point numbers, and booleans. It is mainly used to record information associated with academic papers, covering paper identification numbers, external links, successful contact status, notes, the number of models, the number of datasets, the number of spaces, titles, GitHub-related information, GitHub star counts, conference names, like counts, comment counts, whether GitHub mentions HF, whether there are artifacts, submitters, and dates. The dataset is split into a training set which contains 5064 samples with a total size of 1127665 bytes.
数据集概述
数据集名称
community-science-merged
数据集特征
- arxiv_id: 字符串类型,表示arXiv ID。
- reached_out_link: 字符串类型,表示联系链接。
- reached_out_success: 浮点数类型,表示联系是否成功。
- reached_out_note: 字符串类型,表示联系备注。
- num_models: 浮点数类型,表示模型数量。
- num_datasets: 浮点数类型,表示数据集数量。
- num_spaces: 浮点数类型,表示空间数量。
- title: 字符串类型,表示标题。
- github: 字符串类型,表示GitHub链接。
- github_stars: 浮点数类型,表示GitHub星标数。
- conference_name: 字符串类型,表示会议名称。
- upvotes: 整数类型,表示点赞数。
- num_comments: 整数类型,表示评论数。
- github_mention_hf: 浮点数类型,表示GitHub中提及Hugging Face的情况。
- has_artifact: 布尔类型,表示是否有相关资源。
- submitted_by: 字符串类型,表示提交者。
- date: 字符串类型,表示日期。
数据集分割
- train:
- 字节数: 1128091
- 样本数: 5066
数据集大小
- 下载大小: 392148
- 数据集大小: 1128091
配置文件
- config_name: default
- data_files:
- split: train
- path: data/train-*
- data_files:




