遇见数据集

NeutrinoPit/OpenSubtitles2024-en-ar-batch65

收藏
Hugging Face2025-03-04 更新2025-04-12 收录
官方服务:

资源简介:

该数据集包含两种语言的文本数据:英语(en)和阿拉伯语(ar),均为字符串类型。数据集仅包含训练集部分,共有100万条样本数据,总文件大小约为109MB。数据集提供了默认配置,用于指定训练数据的文件路径。

The dataset includes text data in two languages: English (en) and Arabic (ar), both of which are of string type. The dataset contains only the training set with a total of 1,000,000 samples, with an overall file size of approximately 109MB. The dataset provides a default configuration for specifying the file path of the training data.

提供机构:
NeutrinoPit
二维码
社区交流群
二维码
科研交流群
商业服务