Patents Phrase to Phrase Semantic Matching Dataset
收藏资源简介:
Patents Phrase to Phrase Semantic Matching Dataset 是由谷歌公司创建的一个专注于专利技术概念的语义匹配数据集。该数据集包含近50,000对经过人工评级的短语对,每对短语都附有一个合作专利分类(CPC)作为上下文。数据集通过提取专利中的关键短语并结合上下文CPC分类来创建,旨在解决短语歧义和对抗性关键词匹配问题。此数据集的应用领域主要是在自然语言处理中,特别是在专利和科学出版物的语义文本相似性测量上,以推动模型在处理技术术语方面的性能提升。
The Patents Phrase to Phrase Semantic Matching Dataset is a semantic matching dataset focused on patent technical concepts, developed by Google. It contains nearly 50,000 manually annotated phrase pairs, each paired with a Cooperative Patent Classification (CPC) code as contextual information. The dataset is constructed by extracting key phrases from patents and combining them with their associated CPC classifications, aiming to resolve phrase ambiguity and adversarial keyword matching challenges. Its primary application scenarios lie in natural language processing, particularly for semantic text similarity measurement in patents and scientific publications, to enhance the performance of models in handling technical terminology.

- 1Patents Phrase to Phrase Semantic Matching Dataset谷歌公司 · 2022年



