proteinglm/fold_prediction
收藏资源简介:
Fold Prediction Dataset是一个用于蛋白质序列折叠分类任务的数据集,将蛋白质序列分配到1,195个已知的折叠类别中。该任务的主要应用包括识别新型远程同源蛋白,如新兴的抗生素抗性基因和工业酶。数据集包含三个部分:训练集、验证集和测试集,分别包含12,312、736和3,244个实例。每个实例包含一个蛋白质序列字符串和一个表示折叠类别的整数标签。数据集基于SCOP 1.75版本,发布于2009年。
The Fold Prediction Dataset is used for a scientific classification task, assigning protein sequences to one of 1,195 known fold categories. The primary applications include identifying novel remote homologs in proteins, such as emerging antibiotic-resistant genes and industrial enzymes. The dataset includes a train set (12,312 instances), a valid set (736 instances), and a test set (3,244 instances). Each instance contains a string representing the protein sequence and an integer label indicating which known fold the protein sequence belongs to. The dataset is based on the SCOP 1.75 version, a release from 2009.




