遇见数据集

Learning Rich Representation of Keyphrases from Text

收藏
Mendeley Data2024-03-27 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

In this work, we explore how to learn task-specific language models aimed towards learning rich representation of keyphrases from text documents. We experiment with different masking strategies for pre-training transformer language models (LMs) in discriminative as well as generative settings. In the discriminative setting, we introduce a new pre-training objective - Keyphrase Boundary Infilling with Replacement (KBIR), showing large gains in performance (up to 9.26 points in F1) over SOTA, when LM pre-trained using KBIR is fine-tuned for the task of keyphrase extraction. In the generative setting, we introduce a new pre-training setup for BART - KeyBART, that reproduces the keyphrases related to the input text in the CatSeq format, instead of the denoised original input. This also led to gains in performance (up to 4.33 points in F1@M) over SOTA for keyphrase generation. Additionally, we also fine-tune the pre-trained language models on named entity recognition (NER), question answering (QA), relation extraction (RE), abstractive summarization and achieve comparable performance with that of the SOTA, showing that learning rich representation of keyphrases is indeed beneficial for many other fundamental NLP tasks. As a part of this zip file we release the KBIR model which is continually pre-trained on RoBERTa-Large and also the KeyBART model which is continually pre-trained on BART-Large. Both these models can be used in place of a RoBERTa-Large or BART-Large model in PyTorch codebases and also with HuggingFace.

本研究围绕如何构建面向特定任务的语言模型展开探索,该模型旨在从文本文档中学习关键词短语的丰富表征。我们针对判别式与生成式两种场景下的Transformer语言模型(Transformer)预训练,尝试了多种掩码策略。在判别式场景中,我们提出了一种全新的预训练目标——带替换的关键词短语边界填充(Keyphrase Boundary Infilling with Replacement,KBIR)。当使用KBIR预训练的语言模型针对关键词抽取任务进行微调时,其性能相较当前最优模型(SOTA)实现了大幅提升,F1值最高可提升9.26个百分点。在生成式场景中,我们针对BART模型提出了全新的预训练框架——KeyBART,该框架不再以去噪后的原始输入为生成目标,而是以CatSeq格式生成与输入文本相关的关键词短语。针对关键词生成任务,该模型相较当前最优模型(SOTA)同样实现了性能提升,F1@M值最高可提升4.33个百分点。此外,我们还将预训练语言模型在命名实体识别(Named Entity Recognition,NER)、问答(Question Answering,QA)、关系抽取(Relation Extraction,RE)以及抽象式摘要等任务上进行微调,其性能可与当前最优模型(SOTA)相媲美,这表明学习关键词短语的丰富表征确实对诸多其他基础自然语言处理(Natural Language Processing,NLP)任务具有增益效果。本次发布的压缩包中包含了基于RoBERTa-Large进行持续预训练得到的KBIR模型,以及基于BART-Large进行持续预训练得到的KeyBART模型。这两款模型均可在PyTorch代码库以及HuggingFace平台中替代原生的RoBERTa-Large或BART-Large模型使用。

创建时间:
2023-06-28
二维码
社区交流群
二维码
科研交流群
商业服务