bigbio/umnsrs
收藏资源简介:
--- language: - en bigbio_language: - English license: cc0-1.0 multilinguality: monolingual bigbio_license_shortname: CC0_1p0 pretty_name: UMNSRS homepage: https://conservancy.umn.edu/handle/11299/196265/ bigbio_pubmed: False bigbio_public: True bigbio_tasks: - SEMANTIC_SIMILARITY --- # Dataset Card for UMNSRS ## Dataset Description - **Homepage:** https://conservancy.umn.edu/handle/11299/196265/ - **Pubmed:** False - **Public:** True - **Tasks:** STS UMNSRS, developed by Pakhomov, et al., consists of 725 clinical term pairs whose semantic similarity and relatedness. The similarity and relatedness of each term pair was annotated based on a continuous scale by having the resident touch a bar on a touch sensitive computer screen to indicate the degree of similarity or relatedness. The following subsets are available: - similarity: A set of 566 UMLS concept pairs manually rated for semantic similarity (e.g. whale-dolphin) using a continuous response scale. - relatedness: A set of 588 UMLS concept pairs manually rated for semantic relatedness (e.g. needle-thread) using a continuous response scale. - similarity_mod: Modification of the UMNSRS-Similarity dataset to exclude control samples and those pairs that did not match text in clinical, biomedical and general English corpora. Exact modifications are detailed in the paper (Corpus Domain Effects on Distributional Semantic Modeling of Medical Terms. Serguei V.S. Pakhomov, Greg Finley, Reed McEwan, Yan Wang, and Genevieve B. Melton. Bioinformatics. 2016; 32(23):3635-3644). The resulting dataset contains 449 pairs. - relatedness_mod: Modification of the UMNSRS-Relatedness dataset to exclude control samples and those pairs that did not match text in clinical, biomedical and general English corpora. Exact modifications are detailed in the paper (Corpus Domain Effects on Distributional Semantic Modeling of Medical Terms. Serguei V.S. Pakhomov, Greg Finley, Reed McEwan, Yan Wang, and Genevieve B. Melton. Bioinformatics. 2016; 32(23):3635-3644). The resulting dataset contains 458 pairs. ## Citation Information ``` @inproceedings{pakhomov2010semantic, title={Semantic similarity and relatedness between clinical terms: an experimental study}, author={Pakhomov, Serguei and McInnes, Bridget and Adam, Terrence and Liu, Ying and Pedersen, Ted and Melton, Genevieve B}, booktitle={AMIA annual symposium proceedings}, volume={2010}, pages={572}, year={2010}, organization={American Medical Informatics Association} } ```
--- language: 语言 - 英语 bigbio_language: - 英语 license: 知识共享CC0 1.0通用公共许可协议 multilinguality: 单语 bigbio_license_shortname: CC0_1p0 pretty_name: UMNSRS homepage: https://conservancy.umn.edu/handle/11299/196265/ bigbio_pubmed: 否 bigbio_public: 是 bigbio_tasks: - 语义相似度(SEMANTIC_SIMILARITY) --- # UMNSRS 数据集卡片 ## 数据集描述 - **主页:** https://conservancy.umn.edu/handle/11299/196265/ - **关联PubMed:** 否 - **公开状态:** 是 - **任务:** 语义相似度任务(STS) 由Pakhomov等人研发的UMNSRS数据集包含725对临床术语,所有术语对的语义相似度与相关性均已完成标注。每对术语的相似度与相关性均采用连续量表进行标注:由住院医师通过触摸触控电脑屏幕上的滑块,指示术语对的相似或相关程度。 该数据集提供以下子集: - **相似度子集:** 包含566对**统一医学语言系统(Unified Medical Language System,UMLS)**概念对,采用连续响应量表对其语义相似度进行人工标注(例如:鲸鱼-海豚)。 - **相关性子集:** 包含588对UMLS概念对,采用连续响应量表对其语义相关性进行人工标注(例如:针-线)。 - **修正版相似度子集:** 对原始UMNSRS-相似度数据集进行优化,剔除其中的对照样本以及未在临床、生物医学与通用英语语料库中出现匹配文本的术语对。具体优化细节详见论文《Corpus Domain Effects on Distributional Semantic Modeling of Medical Terms》(作者:Serguei V.S. Pakhomov、Greg Finley、Reed McEwan、Yan Wang、Genevieve B. Melton,发表于《Bioinformatics》,2016年,32(23):3635-3644)。优化后的数据集共包含449对术语对。 - **修正版相关性子集:** 对原始UMNSRS-相关性数据集进行优化,剔除其中的对照样本以及未在临床、生物医学与通用英语语料库中出现匹配文本的术语对。具体优化细节详见上述同一篇论文。优化后的数据集共包含458对术语对。 ## 引用信息 @inproceedings{pakhomov2010semantic, title={Semantic similarity and relatedness between clinical terms: an experimental study}, author={Pakhomov, Serguei and McInnes, Bridget and Adam, Terrence and Liu, Ying and Pedersen, Ted and Melton, Genevieve B}, booktitle={AMIA annual symposium proceedings}, volume={2010}, pages={572}, year={2010}, organization={American Medical Informatics Association} }
数据集概述
名称: UMNSRS
语言: 英语
许可: CC0-1.0
多语言性: 单语
任务: 语义相似性(SEMANTIC_SIMILARITY)
数据集详情
- 开发者: Pakhomov, et al.
- 内容: 包含725个临床术语对,用于评估语义相似性和相关性。
- 评估方法: 通过居民在触摸屏上操作,使用连续响应尺度进行标注。
数据集子集
-
相似性(similarity):
- 数量: 566对UMLS概念对
- 用途: 手动评估语义相似性
-
相关性(relatedness):
- 数量: 588对UMLS概念对
- 用途: 手动评估语义相关性
-
相似性修正(similarity_mod):
- 数量: 449对
- 修改详情: 排除控制样本及未匹配文本的样本,详细修改见相关论文。
-
相关性修正(relatedness_mod):
- 数量: 458对
- 修改详情: 排除控制样本及未匹配文本的样本,详细修改见相关论文。
引用信息
@inproceedings{pakhomov2010semantic, title={Semantic similarity and relatedness between clinical terms: an experimental study}, author={Pakhomov, Serguei and McInnes, Bridget and Adam, Terrence and Liu, Ying and Pedersen, Ted and Melton, Genevieve B}, booktitle={AMIA annual symposium proceedings}, volume={2010}, pages={572}, year={2010}, organization={American Medical Informatics Association} }




