OmniGenBench
收藏资源简介:
OmniGenBench是由埃克塞特大学计算机科学系开发的用于基因组基础模型(GFM)的自动化大规模基准测试框架。该数据集整合了来自四个大规模基准的4200万条基因组序列,涵盖了数百个基因组任务,旨在解决基因组数据稀缺和偏差的问题。数据集的创建过程包括数据过滤和标准化,以确保下游任务的数据质量。OmniGenBench的应用领域广泛,包括基因组序列的合成、RNA结构预测和功能预测等,旨在推动基因组研究的自动化和高效化。
OmniGenBench is an automated large-scale benchmarking framework developed by the Department of Computer Science at the University of Exeter for genomic foundation models (GFM). This dataset integrates 42 million genomic sequences from four large-scale benchmarks, covering hundreds of genomic tasks, aiming to address the issues of genomic data scarcity and bias. The dataset creation process includes data filtering and standardization to ensure data quality for downstream tasks. OmniGenBench has a wide range of application scenarios, including genomic sequence synthesis, RNA structure prediction, functional prediction and more, aiming to promote the automation and efficiency of genomic research.

- 1OmniGenBench: Automating Large-scale in-silico Benchmarking for Genomic Foundation Models埃克塞特大学计算机科学系 · 2024年



