Benchmarking Splits
收藏资源简介:
Benchmarking splits used to evaluate the LinkPhinder and other protein phosphorylation prediction models. The XY numbers in the benchmark folder names reflect the train-test data ratio (e.g. 'benchmark9010' means the training and testing data are 90% and 10%, respectively, of the benchmark dataset). The nZ number then means the number of negatives per each positive (e.g. 'benchmark7030n10' means a 70/30 train/test split with ten negatives per each positive). PhosphoSitePlus (https://doi.org/10.1093/nar/gku1267) was used as the benchmark dataset in most cases, with the exception of folders tagged with 'cutillas20' (this is based on the kinase-substrate interaction dataset published in https://doi.org/10.1038/s41587-019-0391-9). The columns in the training and testing split files correspond to [TODO: fill in the details here].
本基准拆分数据集用于评估LinkPhinder及其他蛋白质磷酸化预测模型。基准文件夹名称中的XY数值代表训练集与测试集的数据占比(例如,"benchmark9010"表示该基准数据集的训练集和测试集分别占总数据的90%与10%)。其中的nZ数值则代表每个正样本所对应的负样本数量(例如,"benchmark7030n10"表示训练测试拆分比为70/30,且每个正样本对应10个负样本)。多数情况下基准数据集采用PhosphoSitePlus(https://doi.org/10.1093/nar/gku1267),仅带有"cutillas20"标记的文件夹除外——该数据集基于发表于https://doi.org/10.1038/s41587-019-0391-9的激酶-底物相互作用数据集。训练与测试拆分文件中的列对应内容为[待补充详细信息]。




