TM-Senti
收藏资源简介:
TM-Senti是由伦敦玛丽女王大学开发的一个大规模、远距离监督的Twitter情感数据集,包含超过1.84亿条推文,覆盖了超过七年的时间跨度。该数据集基于互联网档案馆的公开推文存档,可以完全重新构建,包括推文元数据且无缺失推文。数据集内容丰富,涵盖多种语言,主要用于情感分析和文本分类等任务。创建过程中,研究团队精心筛选了表情符号和表情,确保数据集的质量和多样性。该数据集的应用领域广泛,旨在解决社交媒体情感表达的长期变化问题,特别是在表情符号和表情使用上的趋势分析。
TM-Senti is a large-scale, distant-supervised Twitter sentiment dataset developed by Queen Mary University of London, containing over 184 million tweets spanning more than seven years. Based on the public tweet archive of the Internet Archive, this dataset can be fully reconstructed with complete tweet metadata and no missing tweets. Boasting rich content covering multiple languages, it is primarily utilized for tasks such as sentiment analysis and text classification. During its development, the research team carefully curated emojis and emoticons to guarantee the dataset's quality and diversity. This dataset has a wide range of application scenarios, aiming to address long-term changes in emotional expression on social media, particularly trend analysis of emoji and emoticon usage.

- 1The emojification of sentiment on social media: Collection and analysis of a longitudinal Twitter sentiment dataset伦敦玛丽女王大学 · 2023年



