univer-api-sft
收藏资源简介:
该数据集包含编程相关任务的指令-代码对,共5,685个样本(训练集4,548例,验证集568例,测试集569例)。每条数据包含六个字段:自然语言指令(instruction)、对应代码(code)、任务类别(category)、难度等级(difficulty)、涉及的API方法(api_methods)以及唯一任务ID(task_id)。数据总量约3MB,以纯文本形式存储,按标准机器学习流程划分为训练/验证/测试集。虽然未明确说明应用场景,但数据结构表明其适用于代码生成、程序合成或AI编程助手等NLP任务。
This dataset contains instruction-code pairs for programming-related tasks, with a total of 5,685 samples: 4,548 for the training set, 568 for the validation set, and 569 for the test set. Each sample includes six fields: natural language instruction (`instruction`), corresponding code (`code`), task category (`category`), difficulty level (`difficulty`), involved API methods (`api_methods`), and unique task ID (`task_id`). The total size of the dataset is approximately 3 MB, stored in plain text format, and split into training, validation and test sets following standard machine learning workflows. Although the application scenarios are not explicitly specified, the data structure indicates that it is suitable for NLP tasks such as code generation, program synthesis, or AI programming assistants.




