Sera-4.5A-Full-T1-v2-31600
收藏资源简介:
该数据集包含31,600个训练样本,总大小为4.78GB。每个样本包含三个主要字段:1) conversations字段为对话列表,每条对话包含role(角色)和content(内容)两个字符串字段;2) source字段表示数据来源的字符串;3) instance_id字段为实例标识字符串。数据集采用单一训练集划分,未提供验证集或测试集。数据以分片文件形式存储,路径模式为train-*。
This dataset contains 31,600 training samples with a total size of 4.78 GB. Each sample includes three core fields: 1) The `conversations` field is a list of dialogue turns, where each turn contains two string fields: `role` (speaker role) and `content` (dialogue content); 2) The `source` field is a string representing the data source; 3) The `instance_id` field is a string-type instance identifier. The dataset adopts a single training set split, with no validation set or test set provided. The data is stored in sharded file format, with the path pattern being `train-*`.
数据集概述
基本信息
- 数据集名称: Sera-4.5A-Full-T1-v2-31600
- 发布者/组织: laion
- 数据集地址: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v2-31600
数据规模与结构
- 总样本数: 31,600 条
- 数据分割: 仅包含训练集(train)
- 训练集样本数: 31,600 条
- 训练集数据大小: 4,785,185,963 字节
- 下载文件大小: 1,511,743,433 字节
- 数据集存储大小: 4,785,185,963 字节
数据特征(Features)
数据集包含以下字段:
- conversations(列表类型)
- role: 字符串类型,表示对话中的角色
- content: 字符串类型,表示对话内容
- source: 字符串类型,表示数据来源
- instance_id: 字符串类型,表示实例的唯一标识符
数据文件配置
- 配置名称: default
- 数据文件路径:
data/train-*(对应训练集分割)




