RNAChallenge
收藏资源简介:
该数据集来源于公开数据集,用于通过大规模工具基准测试来分类蛋白质编码和非编码RNA。它是一个具有挑战性的验证数据集,仅包含困难实例,并具有双重目的:1. 作为标准测试数据集,评估工具性能,对于开发更准确、偏差更小的模型以正确注释转录本是必要的。2. 识别转录本误注释问题,通过自动化方法和生物信息学专家识别假阳性和假阴性实例。
This dataset is derived from publicly available datasets and is utilized for classifying protein-coding and non-coding RNAs through large-scale tool benchmarking. It serves as a challenging validation dataset, exclusively comprising difficult instances, and fulfills a dual purpose: 1. As a standard testing dataset, it is essential for evaluating tool performance, which is crucial for developing more accurate and less biased models to correctly annotate transcripts. 2. To identify issues of transcript misannotation, it employs automated methods and bioinformatics experts to discern false positives and false negatives.
RNAChallenge数据集概述
数据集目的
- 标准测试数据集:用于评估工具性能,帮助开发更准确、无偏差的转录本注释模型。
- 识别错误注释:通过自动化方法和生物信息学专家识别假阳性和假阴性实例,解决生物学领域中的转录本错误注释问题。
数据集来源
- 数据收集:包含135个小规模和大规模转录组数据集,覆盖49个物种,涉及动物、植物和真菌界。
- 数据特征:图示展示了物种的界级覆盖、mRNAs和ncRNAs的比例、数据集的平衡性以及序列长度的分布。
数据集构建
- 分类决策:通过平均48个模型的分类决策来获得每个实例的分类置信度分数。
- 数据筛选:过滤掉长度低于特定阈值的序列,并使用CD-HIT工具移除重复序列。
- 数据组成:包含16,243个mRNAs和11,040个ncRNAs,来自动物、植物和真菌界的物种。
引用信息
- 参考文献:Dalwinder Singh和Joy Roy, A large-scale benchmark study of tools for the classification of protein-coding and non-coding RNAs, Nucleic Acids Research, gkac1092, 2022.




