qna
收藏资源简介:
QnA数据集是一个通过多种神经网络模型生成问题答案而创建的多语言问答数据集,包含英语和俄语两种语言,规模中等(样本量级在1K到10K之间)。数据统计显示总token数为12,271,743,其中问题部分占243,559个token,答案部分占11,735,372个token,表明答案内容相对详细丰富。数据集使用了六种不同的神经网络模型进行答案生成,包括DeepSeek-V4-Flash(158B参数)、Gemma-3-1B、Gemma-4-E2B(2.3B参数)、Gemini-3-Flash、Qwen3-4B(4B参数)以及LFM2.5-1.2B-Instruct(1.2B参数)。这些模型生成的问答对适用于自然语言处理任务,如问答系统评估、语言模型训练、答案生成质量比较等。数据集采用CC-BY-4.0许可证发布。
The QnA dataset is a multilingual question-answering dataset created by generating question-answer pairs using multiple neural network models. It includes English and Russian languages and is of medium scale (1K-10K sample range). Data statistics show a total token count of 12,271,743, with 243,559 tokens in the question part and 11,735,372 tokens in the answer part, indicating relatively detailed and rich answer content. The dataset utilizes six different neural network models for answer generation, including DeepSeek-V4-Flash (158B parameters), Gemma-3-1B, Gemma-4-E2B (2.3B parameters), Gemini-3-Flash, Qwen3-4B (4B parameters), and LFM2.5-1.2B-Instruct (1.2B parameters). These generated question-answer pairs are suitable for natural language processing tasks, such as question-answering system evaluation, language model training, and answer generation quality comparison. The dataset is released under the CC-BY-4.0 license.
数据集概述
数据集名称: QnA dataset
许可证: CC-BY-4.0
语言: 英语 (en)、俄语 (ru)
数据规模: 1K < n < 10K(样本数量在1000到10000之间)
数据集内容
该数据集通过多个神经网络模型生成问题对应的答案而创建。
统计信息
| 参数 | 值 |
|---|---|
| 总 Token 数 | 12,271,743 |
| 问题 Token 数 | 243,559 |
| 答案 Token 数 | 11,735,372 |
使用的模型
数据集的问题答案由以下模型生成:
- deepseek-v4-flash (158B 参数) - 模型页面
- google/gemma-3-1b (1B 参数) - 模型页面
- google/gemma-4-e2b (2.3B 参数) - 模型页面
- gemini-3-flash (参数规模未知) - 详细信息
- qwen/qwen3-4b-2507 (4B 参数) - 模型页面
- lfm2.5-1.2b-instruct (1.2B 参数) - 模型页面
数据集用途
该数据集适用于问答任务、多语言文本生成、模型性能对比等自然语言处理相关研究与开发场景。




