Caribou pipeline for the alignment-free bacterial identification and classification in metagenomics sequencing data using machine learning
收藏资源简介:
This dataset contains sequencing data used to train the models of the Caribou pipeline. We developed this pipeline for alignment-free bacterial identification and classification in metagenomics sequencing data using machine learning. The datasets were derived from the GTDB v.202 database (https://data.gtdb.ecogenomic.org/releases/release202/202.0/) and include training steps using the species representatives, as the benchmark datasets used non-representative whole genomes. We also simulated sequencing reads to evaluate and compare performance on whole genomes and sequencing reads. We provide models and encoding files of CNN-trained models; datasets used for training, validation and testing of models, randomly sampled from representative genomes; and datasets used for benchmarking the method against state-of-the-art methods, randomly sampled from non-representative whole genomes and simulated reads.



