emoticon dataset
收藏资源简介:
该数据集名为emoticon dataset,由清华大学DCST和济南量子技术研究院联合创建。这是一个包含10个不同领域、跨语言、时间序列丰富的表情符号用户交互数据集,共包含22K个独特用户,370K个表情符号和8.3M条对话信息。数据集从广泛使用的即时通讯平台中收集,经过严格的数据完整性和安全性检查。该数据集为公开可访问的最大表情符号数据集,可广泛应用于用户行为分析和个性化表情推荐系统等研究。
This dataset, named Emoticon Dataset, was jointly created by the DCST of Tsinghua University and Jinan Institute of Quantum Technology. It is a cross-lingual user interaction dataset rich in temporal sequences, covering 10 distinct domains, with a total of 22K unique users, 370K emojis, and 8.3M conversation messages. The dataset was collected from widely used instant messaging platforms and underwent strict data integrity and security checks. As the largest publicly accessible emoji dataset to date, it can be widely applied to research such as user behavior analysis and personalized emoji recommendation systems.




