遇见数据集

MultiSynt/nemotron-cc-polish-tower9b

收藏
Hugging Face2025-10-15 更新2025-10-25 收录
官方服务:

资源简介:

Nemotron-cc高实际子集,已翻译成波兰语,使用Tower+ 9B模型。数据集包含154,137,284行,共有108,127,852,463个tokens。

Nemotron-cc high actual subset translated to Polish using Tower+ 9B. The dataset contains 154,137,284 rows and consists of 108,127,852,463 tokens.

提供机构:
MultiSynt
二维码
社区交流群
二维码
科研交流群
商业服务