LongABC-32K
收藏资源简介:
LongABC-32K数据集是由北京大学等研究机构创建的高质量长语境训练数据集。该数据集通过开源的长语境数据集(ArXiv、书籍和代码)进行过滤得到,包含了强大的长距离依赖性。数据集的大小为32k tokens,是为了在捕获长距离依赖性和保持合理的计算复杂性之间取得平衡。该数据集的创建是为了增强大型语言模型处理长语境的能力,并已发布以促进未来长语境数据的研究。
The LongABC-32K dataset is a high-quality long-context training dataset created by Peking University and other research institutions. It is filtered from open-source long-context datasets (ArXiv papers, books, and code), and exhibits strong long-range dependencies. With a size of 32k tokens, it is designed to strike a balance between capturing long-range dependencies and maintaining reasonable computational complexity. This dataset was developed to enhance the long-context processing capabilities of large language models, and has been publicly released to facilitate future research on long-context data.

- 1LongAttn: Selecting Long-context Training Data via Token-level Attention北京大学 · 2025年



