T2K2
收藏资源简介:
T2K2是由布加勒斯特理工大学创建的一个包含250万条推文的数据集,旨在通过分析文本数据中的关键词来提取知识。数据集通过Twitter的REST API收集,并经过预处理,包括提取标签、扩展缩写、提取句子和词干等步骤。该数据集被用于评估不同的权重计算方案和数据库实现,特别关注计算性能。T2K2的应用领域包括文本分析、趋势识别和异常检测,旨在通过高效计算关键词来解决信息检索中的问题。
T2K2 is a dataset consisting of 2.5 million tweets created by the University Politehnica of Bucharest, designed to extract knowledge by analyzing keywords in textual data. The dataset was collected via Twitter's REST API and has undergone preprocessing steps including hashtag extraction, abbreviation expansion, sentence segmentation, and stemming. This dataset is used to evaluate different weight calculation schemes and database implementations, with a particular focus on computational performance. The application fields of T2K2 include text analysis, trend identification and anomaly detection, aiming to solve problems in information retrieval through efficient keyword computation.



