MedRedQA
收藏资源简介:
A large non-factoid English consumer Question Answering (QA) dataset containing 51,000 pairs of consumer questions and their corresponding expert answers. This dataset is useful for bench-marking or training systems on more difficult real-world questions and responses which may contain spelling or formatting errors, or lexical gaps between consumer and expert vocabularies.\n\nBy downloading this dataset, you agree to have obtained ethics approval from your institution.\nLineage: We collected data from posts and comments to subreddit /r/askdocs, published between July 10, 2013, and April 2, 2022, totalling 600,000 submissions (original posts) and 1,700,000 comments (replies). We generated question-answer pairs by taking the highest scoring answer from a verified medical expert to a Reddit question. Questions with only images are removed, all links are removed and authors are removed. \n\nWe provide two separate datasets in this collection and provide the following schemas.\nMedRedQA - Reddit Medical Question and Answer pairs from /r/askdocs. CSV format.\ni. the poster's question (Body) \nii. Title of the post \niii. The filtered answer from a verified physician comment (Response)\niv. Occupation indicated for verification status\nv. Any PMCIDs found in the post\n\nMedRedQA+PubMed - PubMed Enriched subset of MedRedQA. JSON format.\ni. Question. The user's original question. The is equivalent to the Body field in MedRedQA\nii. Document: The abstract of the PubMed document (if it exists and contains an abstract) for that particular post. Note: it does not necessarily mean the answer references this document. But at least one other verified physician in the responses has mentioned that particular document.\niii. The filtered response. This is equivalent to the Response field in MedRedQA.
本数据集为大型非事实型英文消费者问答(Question Answering, QA)数据集,包含51000条消费者问题与对应专家答复的配对样本。该数据集可用于基准测试,或针对更具挑战性的真实世界问答场景训练系统——这类场景中的问题与答复可能存在拼写、格式错误,或消费者与专家词汇间的语义鸿沟。 下载本数据集即代表您同意已获得所在机构的伦理审查批准。 数据集溯源:我们从2013年7月10日至2022年4月2日期间发布的Reddit子版块/r/askdocs的帖子与评论中采集数据,总计包含60万条投稿(原始帖子)与170万条评论(回复)。我们通过选取针对Reddit平台上某一问题的、经认证的医学专家所获最高赞答复,来构建问答配对样本。仅包含图片的问题已被移除,所有链接与作者信息均已完成脱敏处理。 本数据集集合包含两个独立子数据集,其结构规范如下: MedRedQA:源自/r/askdocs的Reddit医学问答配对数据集,格式为逗号分隔值(CSV)。 i. 发帖者的问题(Body) ii. 帖子标题(Title) iii. 经认证的医师评论中筛选出的答复(Response) iv. 用于验证认证状态的职业信息 v. 帖子中提及的所有PubMed中心标识符(PMCID) MedRedQA+PubMed:MedRedQA的PubMed增强子集,格式为JavaScript对象表示法(JSON)。 i. 问题(Question):用户的原始问题,与MedRedQA中的Body字段含义完全一致。 ii. 文献(Document):对应帖子关联的PubMed文献摘要(若该文献存在且包含摘要内容)。注意:这并不代表答复参考了该文献,但至少有一位其他经认证的医师在回复中提及了该文献。 iii. 筛选后的答复,与MedRedQA中的Response字段含义完全一致。




