Synthetic dataset for Recommender Systems
收藏资源简介:
本研究聚焦于为推荐系统生成合成数据集,采用统计抽样方法、多项逻辑模型和模糊推理系统。数据集旨在解决推荐系统领域因缺乏高质量数据而面临的挑战,特别是在旅游行业中。该数据集包含数值/序数和名义特征,通过高斯相依函数、狄利克雷和高斯分布、多项逻辑模型以及模糊逻辑推理系统生成评分,根据不同的用户行为模式和感知物品质量进行调整。数据集的创建过程涉及多种技术,包括用户特征、物品属性和类别以及潜在用户偏好的定义,最终形成用户-物品稀疏评分矩阵。该数据集应用于旅游推荐系统,旨在通过模拟真实用户行为和偏好,提高推荐算法的准确性和实用性。
This study focuses on generating synthetic datasets for recommendation systems, leveraging statistical sampling methods, multinomial logistic models, and fuzzy inference systems. This dataset aims to address the challenges faced by the recommendation system domain due to the shortage of high-quality data, particularly in the tourism industry. It contains numerical/ordinal and nominal features, with ratings generated via Gaussian copula functions, Dirichlet and Gaussian distributions, multinomial logistic models, and fuzzy logic inference systems, and adjusted based on diverse user behavior patterns and perceived item quality. The dataset creation process involves multiple techniques, including the definition of user characteristics, item attributes and categories, as well as latent user preferences, ultimately forming a user-item sparse rating matrix. This dataset is applied to tourism recommendation systems, with the goal of enhancing the accuracy and practicality of recommendation algorithms by simulating real user behaviors and preferences.




