GRIMM
收藏资源简介:
GRIMM是由佐治亚大学团队开发的酶功能预测基准数据集,基于UniProt/SwissProt数据库中的原核生物氨基酸序列构建,包含185,418条训练序列和52,003条测试序列(含开放集测试)。该数据集通过遗传分层技术(UniRef50聚类)确保训练与测试集间的序列相似性隔离,并划分封闭集(Test-1)和开放集(Test-2)以评估模型对新酶功能的发现能力。其创新性在于将生物建模中隐式的序列聚类策略标准化,为蛋白质功能预测等任务提供可复现的评估框架。
GRIMM is a benchmark dataset for enzyme function prediction developed by the team at the University of Georgia. It is constructed based on prokaryotic amino acid sequences from the UniProt/SwissProt database, containing 185,418 training sequences and 52,003 test sequences including open-set testing. This dataset adopts UniRef50 clustering as a genetic stratification technique to ensure the isolation of sequence similarity between the training and test sets, and partitions closed-set (Test-1) and open-set (Test-2) subsets to evaluate the model's ability to discover novel enzyme functions. Its innovation lies in standardizing the implicit sequence clustering strategies in biological modeling, providing a reproducible evaluation framework for tasks such as protein function prediction.
- 1GRIMM: Genetic stRatification for Inference in Molecular Modeling佐治亚大学·海洋科学系 · 2026年



