kgrabko/JiRackDatasetWikipediaEn
收藏资源简介:
该数据集专为JiRack分词器格式化,用于训练JiRack Base 1.5B模型。建议初始化模型时使用4K上下文窗口以确保稳定性,随后通过专门的JiRack 8K数据集扩展到8K上下文,这种两阶段方法在扩展模型长距离依赖之前确保稳健的位置编码。数据集针对银行和金融科技机构设计,用于构建安全、内部的模型,支持端到端解决方案,以预训练模型用于欺诈预防、垃圾邮件过滤、风险评估和反洗钱检测。这是基础检查点,在领域特定数据集微调前进行评估,主要目标是验证初始预训练阶段后RoPE(旋转位置嵌入)的稳定性和一致性。数据集基于广泛的11亿令牌语料库训练,参数与令牌比例接近7:1,在轻量级架构中实现高知识密度和推理能力。性能方面,JiRack Ternary Pro 1.5B在NVIDIA BlackWell 96 Gb VRAM上每轮训练约需14-18小时。数据集仅限个人和非商业研究使用,商业用途需获得许可并支付5%的版税。
The dataset is formatted for the JiRack tokenizer for the JiRack Base 1.5B model. It is recommended to initialize the model with a 4K context window for initial stability, followed by scaling to 8K context using specialized JiRack 8K datasets. This two-stage approach ensures robust positional encoding before extending the models long-range dependency. Designed for banking and fintech institutions, it enables building secure, internal models tailored for the banking sector, providing end-to-end solutions to pre-train models for fraud prevention, spam filtering, risk assessment, and Anti-Money Laundering (AML) detection. This is the base checkpoint, evaluated prior to fine-tuning on domain-specific datasets, with the primary objective of validating RoPE (Rotary Positional Embeddings) stability and coherence following the initial pre-training phase. The dataset is trained on an extensive 11 billion token corpus with a token-to-parameter ratio of nearly 7:1, achieving exceptional knowledge density and reasoning capabilities in a lightweight architecture. Performance-wise, JiRack Ternary Pro 1.5B takes about 14–18 hours per epoch on NVIDIA BlackWell 96 Gb VRAM. It is allowed for personal and non-commercial research use only, with commercial use requiring a license and a 5% royalty on net revenue.



