遇见数据集

Codon seqence predictions using Hyna DNA and Prot Mamba

收藏
Zenodo2026-05-12 更新2026-05-26 收录
官方服务:

资源简介:

The primary focus of this project is codon sequence prediction, given the protein sequence and the DNA surrounding the gene. The dataset includes approximately 42,000 training samples, with the validation and test sets containing roughly 5,000 samples each. To prevent data leakage, the split was performed using homology clusters, ensuring that homologous genes remain within the same group. Protein abundance (PA) data was extracted from PaxDb, while DNA, protein, and codon sequences were sourced from NCBI. PA values were categorized into six classes based on abundance levels. These categories, along with organism tokens (1–17), were utilized as auxiliary tasks for the prediction model.

提供机构:
Zenodo
创建时间:
2026-05-12
二维码
社区交流群
二维码
科研交流群
商业服务