关键短语抽取
收藏资源简介:
针对某时间段内反应民情的工单数据,该能力可自动抽取其中的热点问题,传统的分词+统计的方法需要大量人工标注关键词。本能力采用基于信息熵和内部凝固度的词库构建算法,可方便快捷的实现无监督的词库建立,再通过语言模型的文本表征能力计算各个词或短语对于句子的重要性,从而抽取其中的关键词
Targeting work order data reflecting public sentiment over a given time period, this capability enables automatic extraction of hot issues from such datasets. Traditional word segmentation plus statistical approaches require extensive manual keyword annotation. This capability adopts a vocabulary construction algorithm based on information entropy and internal cohesion, allowing for fast and convenient unsupervised vocabulary construction. Subsequently, it leverages the text representation capability of language models to calculate the importance of each word or phrase to the corresponding sentence, thereby extracting the core keywords from the data.




