遇见数据集

nuvocare/MSD_instruct

收藏
Hugging Face2024-03-10 更新2024-06-22 收录
官方服务:

资源简介:

--- dataset_info: features: - name: User dtype: string - name: Category dtype: string - name: Language dtype: class_label: names: '0': english '1': french '2': german '3': spanish - name: Topic1 dtype: string - name: Topic2 dtype: string - name: Topic3 dtype: string - name: Text dtype: string - name: Question dtype: string splits: - name: train num_bytes: 133987453 num_examples: 79898 - name: test num_bytes: 44598046 num_examples: 26639 download_size: 107452363 dataset_size: 178585499 configs: - config_name: default data_files: - split: train path: data/train-* - split: test path: data/test-* license: apache-2.0 task_categories: - text-generation - text2text-generation language: - de - es - fr - en tags: - medical size_categories: - 10K<n<100K --- # MSD_manual_topics_user_base This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience. The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version. The content, while being labelled the same, differs by the type of user in order to facilitate understanding for patients or give clear details for professional. The manual is available in different languages. This dataset focuses on spanish, german, english and french content about health topics and symptoms. The content is tagged by 2 to 3 medical topics and flagged by user's type and languages. It consists of roughly 21M words representing 45M tokens. This dataset is built for instruction fine-tuning. We built the "Question" by querying a vanilla Mistral 7B model with the following prompt: ```python You will be asked to create one or several questions in the appropriate language based on three elements. Return the ouptuts in the format of the examples. If asked several, splits the answers with a "&" sign. Example input: For question 1 : elements are musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis and language is english Example output: ["Question 1", "musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis", "English", "How to diagnose a autoimmune Myositis ? "] Example input: For question 514 : elements are troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction and language is french Example output: ["Question 514", "troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction", "French", "Donne moi des informations introductives sur le bloc auriculoventriculaire."] Example input: For question 514 : elements are troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction and language is french For question 1 : elements are musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis and language is english Example output: ["Question 514", "troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction", "French", "Donne moi des informations introductives sur le bloc auriculoventriculaire."] & ["Question 1", "musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis", "English", "How to diagnose a autoimmune Myositis ? "] [/INST] ``` This dataset can be used to fine-tune a model to a task of supproting patients and clinicians to be better informed in an adapted manner. An instruct-free version is available here : https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base This dataset is built using the website : https://www.msdmanuals.com/ provided by Merck & Co. All credits of the contents are for the MSD organization. [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)

dataset_info: features: - name: 用户(User) dtype: 字符串(string) - name: 类别(Category) dtype: 字符串(string) - name: 语言(Language) dtype: class_label: names: '0': 英语(english) '1': 法语(french) '2': 德语(german) '3': 西班牙语(spanish) - name: 主题1(Topic1) dtype: 字符串(string) - name: 主题2(Topic2) dtype: 字符串(string) - name: 主题3(Topic3) dtype: 字符串(string) - name: 文本(Text) dtype: 字符串(string) - name: 问题(Question) dtype: 字符串(string) splits: - name: 训练集(train) num_bytes: 133987453 num_examples: 79898 - name: 测试集(test) num_bytes: 44598046 num_examples: 26639 download_size: 107452363 dataset_size: 178585499 configs: - config_name: 默认(default) data_files: - split: 训练集(train) path: data/train-* - split: 测试集(test) path: data/test-* license: Apache 2.0 task_categories: - 文本生成(text-generation) - 文本到文本生成(text2text-generation) language: - 德语(de) - 西班牙语(es) - 法语(fr) - 英语(en) tags: - 医疗(medical) size_categories: - 10K<n<100K # MSD_manual_topics_user_base 本数据集基于默克公司(Merck & Co)面向大众推出的官方网站https://www.msdmanuals.com/构建。 《默克诊疗手册(MSD Manual)》是覆盖症状、疾病、健康及相关领域的权威知识来源。为兼顾专业医护人员与普通患者的阅读需求,该手册推出了两个独立版本:尽管主题标注一致,但会根据目标用户调整内容深浅——面向患者的版本力求通俗易懂,面向专业人士的版本则提供详尽细节。本手册支持多语言版本。 本数据集聚焦于西班牙语、德语、英语、法语四类语言的健康主题与症状相关内容,每条数据均标注2至3个医学主题,并附带用户类型与语言标签。 数据集总词量约2100万,对应4500万个Token。 本数据集专为大语言模型(Large Language Model,LLM)指令微调设计。其中的“问题(Question)”字段通过以下提示词调用基础版Mistral 7B模型生成: python You will be asked to create one or several questions in the appropriate language based on three elements. Return the ouptuts in the format of the examples. If asked several, splits the answers with a "&" sign. Example input: For question 1 : elements are musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis and language is english Example output: ["Question 1", "musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis", "English", "How to diagnose a autoimmune Myositis ? "] Example input: For question 514 : elements are troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction and language is french Example output: ["Question 514", "troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction", "French", "Donne moi des informations introductives sur le bloc auriculoventriculaire."] Example input: For question 514 : elements are troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction and language is french For question 1 : elements are musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis and language is english Example output: ["Question 514", "troubles cardiaques et vasculaires, Bloc auriculoventriculaire and Introduction", "French", "Donne moi des informations introductives sur le bloc auriculoventriculaire."] & ["Question 1", "musculoskeletal and connective tissue disorders, Autoimmune Myositis and Diagnosis of Autoimmune Myositis", "English", "How to diagnose a autoimmune Myositis ? "] [/INST] 本数据集可用于微调模型,以辅助患者与临床医师获取适配的健康信息。 无指令版本可在此获取:https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base 本数据集仍基于https://www.msdmanuals.com/构建,所有内容版权归默沙东(MSD)组织所有。 [更多信息需求](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)

提供机构:
nuvocare
原始信息汇总

数据集概述

数据集信息

特征

  • User: 字符串类型
  • Category: 字符串类型
  • Language: 分类标签类型,包含以下类别:
    • 0: 英语
    • 1: 法语
    • 2: 德语
    • 3: 西班牙语
  • Topic1: 字符串类型
  • Topic2: 字符串类型
  • Topic3: 字符串类型
  • Text: 字符串类型
  • Question: 字符串类型

数据分割

  • train:
    • 字节数: 133987453
    • 样本数: 79898
  • test:
    • 字节数: 44598046
    • 样本数: 26639

数据集大小

  • 下载大小: 107452363 字节
  • 数据集大小: 178585499 字节

配置

  • default:
    • 训练数据文件路径: data/train-*
    • 测试数据文件路径: data/test-*

许可证

  • apache-2.0

任务类别

  • 文本生成
  • 文本到文本生成

语言

  • 德语
  • 西班牙语
  • 法语
  • 英语

标签

  • 医疗

数据集大小类别

  • 10K<n<100K
搜集汇总
数据集介绍
nuvocare/MSD_instruct 数据集图片
构建方式
本数据集源自默克公司旗下MSD诊疗手册官方网站(https://www.msdmanuals.com/),该手册是涵盖症状、疾病及健康领域的权威知识库,并针对专业人士与普通患者分别提供不同版本的内容。数据集聚焦于西班牙语、德语、法语和英语四种语言的健康主题与症状描述,每条样本包含用户类型、语言标签、三级医学主题分类(Topic1至Topic3)、原始文本及自动生成的问题。问题的构建采用指令微调范式,通过向vanilla Mistral 7B模型输入预设提示模板,要求其基于主题层级与语言信息生成对应语言的问题,最终形成约79,898条训练样本与26,639条测试样本,总词量约2,100万词,对应4,500万token。
特点
该数据集的核心特色在于其多维度结构化设计。首先,内容按用户类型(专业人士/患者)差异化呈现,确保信息传递的适配性;其次,采用三级医学主题标签(Topic1至Topic3)实现细粒度知识组织,覆盖从系统性疾病到具体诊疗环节的层级关系。语言覆盖四种主要欧洲语言,支持跨语言医学问答研究。数据规模适中(10万以内样本),但文本密度高,平均每条样本包含丰富医学信息。此外,问题由大语言模型基于主题与语言自动生成,模拟真实临床咨询场景,兼具专业性与多样性。
使用方法
该数据集专为指令微调任务设计,适用于文本生成与文本到文本生成场景。使用时可直接加载HuggingFace中的'train'与'test'分片,每条数据包含'User'(用户类型)、'Category'(分类)、'Language'(语言)、'Topic1-3'(医学主题层级)、'Text'(原始文本)及'Question'(生成问题)字段。研究者可基于'Text'与'Question'构建输入-输出对,微调模型以提升其在多语言、多用户场景下的医学问答能力。也可将'Topic'层级作为条件输入,引导模型生成特定知识维度的回答。无指令版本数据集(nuvocare/MSD_manual_topics_user_base)可供对比实验。
背景与挑战
背景概述
在医疗健康领域,精准且易于理解的多语言医学知识库对于提升患者与专业人员的沟通效率至关重要。nuvocare/MSD_instruct数据集由nuvocare团队基于默克公司(Merck & Co.)运营的MSD手册网站构建,创建时间约为2023年,核心研究问题在于如何通过指令微调使大语言模型能够根据用户类型(患者或专业人士)和语言(英语、法语、德语、西班牙语)生成适配的医学问答。该数据集覆盖约10万条样本,包含症状、疾病等主题的标注内容,并利用Mistral 7B模型生成问题,为医学领域自然语言处理提供了高质量的指令微调资源,对推动多语言、多受众的智能医疗助手发展具有显著影响力。
当前挑战
该数据集面临的核心挑战包括:其一,医学领域问题的复杂性要求模型精准区分专业与通俗表述,例如同一疾病需为患者生成简洁解释,为医生提供详细诊断细节,这考验了指令微调对用户类型标签的泛化能力。其二,构建过程中,从MSD手册爬取内容需处理多语言对齐与主题标签一致性,而利用Mistral 7B自动生成问题时,模型可能引入语义偏差或错误,例如对罕见疾病的提问不够准确。此外,数据集规模(约45M tokens)在覆盖广泛医学主题时仍可能遗漏边缘案例,且版权归属默克公司需确保使用合规,这些因素均增加了数据质量控制的难度。
常用场景
经典使用场景
MSD_instruct数据集以默沙东诊疗手册这一权威医学知识库为基石,精心构建了面向多语种、多用户类型的指令微调语料。其经典使用场景在于为大型语言模型提供高质量的医学领域指令数据,涵盖英语、法语、德语和西班牙语四种语言,并区分专业医师与普通患者两类受众。数据集中的每条样本均由医学文本、用户类型标签及基于内容生成的提问构成,特别适用于训练模型在医疗问答、症状解释、疾病科普等任务中生成精准且适配受众的回复。这种设计使得模型能够掌握从专业术语到通俗表达的灵活转换能力,成为医疗领域对话系统开发的核心资源。
解决学术问题
该数据集有效回应了医学自然语言处理中两个关键学术挑战:多语言医疗知识的统一表征与用户画像驱动的信息适配。传统医学数据集多聚焦单一语言或单一受众,难以满足全球化背景下跨语言医疗信息服务的需求。MSD_instruct通过引入语言与用户类型双维度标签,使研究者能够探索模型在跨语言知识迁移中的表现,同时解决面向不同认知水平用户的表达风格自适应问题。其意义在于推动了医疗AI从通用对话向精准健康传播的范式转变,为构建更公平、可及的数字化医疗体系提供了数据基础,尤其对低资源语言的医学NLP研究具有重要推动作用。
衍生相关工作
基于MSD_instruct已衍生出多项开创性研究。例如,有工作利用其多语言特性探索跨语言医学问答的零样本迁移能力,验证了在英语数据上训练的模型可泛化至西班牙语或法语场景。另有研究聚焦用户类型感知的文本生成,通过对比专业版与患者版数据的微调效果,提出层级式知识蒸馏框架,使单一模型同时兼顾精准性与可读性。此外,该数据集被用于构建医学领域的检索增强生成(RAG)系统,将MSD手册作为外部知识库,结合指令微调提升模型对罕见病或复杂病例的应答可靠性。这些工作共同拓展了医学NLP在少样本学习、多任务学习等方向的研究边界。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务