Codon seqence predictions using Hyna DNA and Prot Mamba
收藏官方服务:
资源简介:
The primary focus of this project is codon sequence prediction, given the protein sequence and the DNA surrounding the gene. The dataset includes approximately 42,000 training samples, with the validation and test sets containing roughly 5,000 samples each. To prevent data leakage, the split was performed using homology clusters, ensuring that homologous genes remain within the same group. Protein abundance (PA) data was extracted from PaxDb, while DNA, protein, and codon sequences were sourced from NCBI. PA values were categorized into six classes based on abundance levels. These categories, along with organism tokens (1–17), were utilized as auxiliary tasks for the prediction model.
提供机构:
Zenodo创建时间:
2026-05-12



