bri25yu/flores200_baseline_all_mt5
收藏资源简介:
--- dataset_info: features: - name: id dtype: int32 - name: input_ids sequence: int32 - name: attention_mask sequence: int8 - name: labels sequence: int64 splits: - name: train num_bytes: 30474393243 num_examples: 41287764 - name: val num_bytes: 3791994 num_examples: 5000 - name: test num_bytes: 7604613 num_examples: 10000 download_size: 15127185362 dataset_size: 30485789850 --- # Dataset Card for "flores200_baseline_all_mt5" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征字段: - 名称:id,数据类型:32位整数 - 名称:输入标识序列(input_ids),数据类型:32位整数序列 - 名称:注意力掩码(attention_mask),数据类型:8位整数序列 - 名称:标签序列(labels),数据类型:64位整数序列 数据集划分: - 划分名称:训练集(train),字节大小:30474393243,样本数量:41287764 - 划分名称:验证集(val),字节大小:3791994,样本数量:5000 - 划分名称:测试集(test),字节大小:7604613,样本数量:10000 下载总大小:15127185362 数据集总存储大小:30485789850 # 「flores200_baseline_all_mt5」数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
数据集名称
flores200_baseline_all_mt5
数据集特征
- id:整数类型,数据类型为 int32。
- input_ids:序列类型,数据类型为 int32。
- attention_mask:序列类型,数据类型为 int8。
- labels:序列类型,数据类型为 int64。
数据集划分
- 训练集:包含 41287764 个样本,占用空间 30474393243 字节。
- 验证集:包含 5000 个样本,占用空间 3791994 字节。
- 测试集:包含 10000 个样本,占用空间 7604613 字节。
数据集大小
- 下载大小:15127185362 字节。
- 数据集总大小:30485789850 字节。



