CHBench
收藏资源简介:
CHBench是由吉林大学开发的首个全面的中文健康相关基准数据集,旨在评估大型语言模型(LLMs)在理解和生成健康相关信息方面的能力。该数据集包含6493条心理健康相关数据和2999条生理健康相关数据,涵盖广泛的主题。数据来源于网络帖子、考试和现有数据集,经过精心筛选和标注,确保数据的质量和多样性。CHBench的创建过程包括数据收集、黄金标准响应选择和提示-响应对的标注,旨在为评估中文LLMs在健康领域的性能提供基础。该数据集的应用领域主要集中在健康信息的准确性和可靠性评估,旨在解决大型语言模型在健康相关查询中可能存在的误解和错误信息传播问题。
CHBench is the first comprehensive Chinese health-related benchmark dataset developed by Jilin University, designed to evaluate the capabilities of large language models (LLMs) in understanding and generating health-related information. This dataset includes 6,493 mental health-related entries and 2,999 physical health-related entries, covering a broad spectrum of topics. The data is sourced from online posts, examinations and existing datasets, and has been rigorously screened and annotated to ensure its quality and diversity. The development pipeline of CHBench encompasses three core steps: data collection, gold-standard response selection, and annotation of prompt-response pairs, aiming to provide a foundational resource for evaluating the performance of Chinese LLMs in the healthcare domain. The primary applications of this dataset focus on assessing the accuracy and reliability of health information, with the goal of addressing the problems of misinterpretation and misinformation dissemination that LLMs may face in response to health-related queries.
CHBench 数据集概述
数据集简介
CHBench 是一个综合性的中文健康相关基准数据集,旨在评估大型语言模型(LLMs)在理解和处理各种场景下的身心健康问题的能力。该数据集包含 6,493 条与心理健康相关的条目和 2,999 条与身体健康相关的条目,涵盖了广泛的主题。
数据集组成
- 心理健康条目:6,493 条
- 身体健康条目:2,999 条
数据收集与评估
数据集的收集步骤如下:
- 使用 5 个中文语言模型生成响应,并对这些响应进行评估。
- 评估的语言模型包括:ERNIE Bot、Qwen、Baichuan、ChatGLM 和 SparkDesk。
关键发现
- ERNIE Bot 在大多数提示下提供了最佳的整体响应,因此被用作黄金标准响应。
- 敏感问题被排除在外,因为 ERNIE Bot 未能为这些问题生成有效的响应。
- 最终的 CHBench 语料库:2,999 条身体健康条目,6,493 条心理健康条目。
相似性分析
身体健康相似性分析
- ChatGLM 在相似性方面表现最佳,与黄金标准响应的相似度最高。
- Qwen 在某些查询中标记为有毒,但在高相似性范围内表现良好,但产生了很多无效输出。
- SparkDesk 的表现一般。
- Baichuan 通过给出中性响应来避免有毒查询的错误,导致更多数据分布在低和中等相似性区间。
心理健康相似性分析
- SparkDesk 在高相似性范围内表现最佳,但对某些公共帖子和缩写缺乏理解。
- ChatGLM 和 Qwen 也表现良好,但在中等相似性范围内有更多响应,表明存在一定的不一致性。
- Qwen 对数据更为敏感,经常将内容标记为有毒。
- Baichuan 由于频繁的无效输出,分布更为均匀。
注意事项
- 数据集中可能包含被认为具有冒犯性的模型输出。

- 1CHBench: A Chinese Dataset for Evaluating Health in Large Language Models吉林大学 · 2024年



