政务公开信息数据集
收藏资源简介:
为构建政务大语言模型和领域知识库,本文从各官方信息发布网站中收集并整理了一个包含文档和问答对数据的综合性政务公开信息数据集.数据集中的大部分数据源自各政府门户网站及多个政务信息公开平台,部分问答对数据由 ChatGPT 3.5生成,并经人工筛选精炼得到.政务公开信息数据集包含 1900 篇公开政务相关文档和 10503 条问答对.
To build government affairs large language models and domain knowledge bases, this paper collects and organizes a comprehensive government public information dataset comprising document and question-answer pair data from various official information publishing websites. Most of the data in the dataset is sourced from various government portals and multiple government information disclosure platforms, while some question-answer pairs are generated by ChatGPT 3.5 and refined via manual screening. This government public information dataset contains 1900 public documents related to government affairs and 10503 question-answer pairs.




