soketlabs/bhasha-wiki-indic
收藏资源简介:
该数据集包含与印度相关的维基百科文章,并翻译成六种印度语言(印地语、孟加拉语、古吉拉特语、泰米尔语、卡纳达语和乌尔都语)。数据集由Soket AI Labs策划,旨在为需要印度知识和上下文理解的预训练语言模型提供支持。每篇文章包含ID、URL、标题、文本内容以及句子、字符、单词和标记的数量。数据集的创建过程包括从维基百科中筛选与印度相关的文章,清理数据,并使用AI4Bharat的IndicTrans2模型将文章翻译成六种印度语言。
This dataset contains Wikipedia articles related to India, translated into six Indian languages: Hindi, Bengali, Gujarati, Tamil, Kannada, and Urdu. Curated by Soket AI Labs, this dataset is designed to support pretrained language models that require Indian knowledge and contextual understanding. Each article includes an ID, URL, title, text content, along with the counts of sentences, characters, words, and tokens. The dataset creation process involves filtering India-associated articles from Wikipedia, conducting data cleaning, and translating the articles into the six target Indian languages using AI4Bharat's IndicTrans2 model.
数据集概述
基本信息
- 语言: 孟加拉语 (bn), 英语 (en), 古吉拉特语 (gu), 印地语 (hi), 卡纳达语 (kn), 泰米尔语 (ta), 乌尔都语 (ur)
- 许可证: cc-by-3.0
- 大小类别: 1M<n<10M
- 任务类别: 文本生成, 填充掩码
- 任务ID: 语言建模, 掩码语言建模
配置详情
20231101.bn
- 数据文件:
- 分割: 训练
- 路径: ben_Beng/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 674539757
- 样本数: 200820
- 训练:
- 下载大小: 652782434
- 数据集大小: 652782434
20231101.en
- 数据文件:
- 分割: 训练
- 路径: eng_Latn/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 703955598
- 样本数: 200820
- 训练:
- 下载大小: 426488108
- 数据集大小: 426488108
20231101.gu
- 数据文件:
- 分割: 训练
- 路径: guj_Gujr/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 668666407
- 样本数: 200820
- 训练:
- 下载大小: 658661502
- 数据集大小: 658661502
20231101.hi
- 数据文件:
- 分割: 训练
- 路径: hin_Deva/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 678769726
- 样本数: 200820
- 训练:
- 下载大小: 640983312
- 数据集大小: 640983312
20231101.kn
- 数据文件:
- 分割: 训练
- 路径: kan_Knda/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 708769566
- 样本数: 200820
- 训练:
- 下载大小: 689888426
- 数据集大小: 689888426
20231101.ta
- 数据文件:
- 分割: 训练
- 路径: tam_Taml/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 781041863
- 样本数: 200820
- 训练:
- 下载大小: 721062888
- 数据集大小: 721062888
20231101.ur
- 数据文件:
- 分割: 训练
- 路径: urd_Arab/train-*
- 特征:
- id: 字符串
- url: 字符串
- title: 字符串
- text: 字符串
- sents: 整数32
- chars: 整数32
- words: 整数32
- tokens: 整数32
- 分割:
- 训练:
- 字节数: 655510379
- 样本数: 200820
- 训练:
- 下载大小: 543259766
- 数据集大小: 543259766




