Universal Synthetic Dataset for Machine Learning on Spectroscopic Data
收藏资源简介:
本研究创建了一个名为‘Universal Synthetic Dataset for Machine Learning on Spectroscopic Data’的数据集,由卡尔斯鲁厄理工学院自动化与应用信息研究所开发。该数据集包含35,000条人工合成的光谱数据,模拟了X射线衍射、核磁共振和拉曼光谱等多种实验测量技术。数据集的创建过程允许用户根据具体问题调整扫描长度和峰值计数等参数。此数据集主要用于机器学习模型的验证,特别是在光谱数据的自动分类领域,旨在通过模拟数据提高分类任务的性能。
This study presents a newly developed dataset titled "Universal Synthetic Dataset for Machine Learning on Spectroscopic Data", constructed by the Institute of Automation and Applied Informatics at Karlsruhe Institute of Technology (KIT). This dataset comprises 35,000 synthetic spectroscopic data samples that emulate a range of experimental measurement techniques including X-ray diffraction (XRD), nuclear magnetic resonance (NMR), and Raman spectroscopy. The dataset's construction pipeline enables users to tailor key parameters such as scan length and peak count to suit specific research objectives. This dataset is primarily intended for validating machine learning models, with a particular focus on the automated classification of spectroscopic data, aiming to enhance the performance of classification tasks using synthetic data.




