toy-multistep-nn_10-na_20-nab_20-seed_2
收藏资源简介:
该数据集包含四个字段:提示(prompts)、完成(completions)、遮蔽数量(num_maskeds)和文本(texts)。其中,提示和完成字段是字符串类型,用于存储文本提示和相应的完成文本;遮蔽数量字段是整型,用于存储遮蔽的单词数量;文本字段是字符串类型,可能包含原始文本数据。数据集分为训练集(train)、测试集(test_rl)和另一个测试集(test),每个集合都包含262144个示例。数据集的总下载大小为33865975字节,总数据大小为79432884字节。
This dataset contains four fields: prompts, completions, num_maskeds, and texts. The prompts and completions fields are string-type fields used to store text prompts and their corresponding completion texts; the num_maskeds field is an integer-type field used to store the number of masked words; the texts field is a string-type field that may contain raw text data. The dataset is divided into three subsets: the training set (train), the test set (test_rl), and another test set (test), with each subset containing 262144 examples. The total download size of the dataset is 33865975 bytes, and the total data size is 79432884 bytes.




