ailab-bio/TACK
收藏资源简介:
TACK(TArgeting Chimeras Knowledge)是一个用于PROTAC(蛋白水解靶向嵌合体)降解活性预测的、经过整理且适合机器学习的数据集和基准。该数据集从三个主要公共存储库(TPDdb、PROTAC-DB、PROTACpedia)中收集并标准化了数据,解决了现有基准中的数据稀缺、不一致性和范围有限的问题。它包含6,561个降解端点,涵盖4,184个DC50(效力)测量值和2,377个Dmax(最大降解功效)测量值,涉及3,514个独特PROTACs、164个POI(目标蛋白)、9个E3连接酶和155个细胞系。数据集提供四个配置:default(所有端点组合)、DC50(仅效力测量)、Dmax(仅最大功效测量)和multitask(同一PROTAC/测定中配对的DC50和Dmax数据,用于二元活性分类)。每个配置包含丰富的特征,如SMILES(PROTAC的规范表示)、POI和连接酶的序列与UniProt注释、细胞系信息(基于Cellosaurus标准化)、测量值(包括单位、误差、范围等)、实验条件(如测定时间、浓度)、来源数据库引用以及用于交叉验证和保持集的结构聚类标识。数据集还包含一个结构上不同的保持集(约10%数据),用于严格的基准评估,并支持回归(pDC50、Dmax)和二元分类任务。基准测试结果显示,在保持集上,pDC50的预测性能(R²为0.66)优于Dmax(R²为0.36),且经典树基方法(如XGBoost)在分类任务中优于特定领域GNN。数据集旨在促进PROTAC设计的机器学习研究,并提供预训练模型和统计评估框架。
TACK (TArgeting Chimeras Knowledge) is a curated, machine learning-ready dataset and benchmark for PROTAC (Proteolysis-Targeting Chimera) degradation activity prediction. It aggregates and standardizes data from three major public repositories (TPDdb, PROTAC-DB, PROTACpedia), addressing gaps in existing benchmarks such as data scarcity, inconsistency, and limited scope. The dataset contains 6,561 degradation endpoints, including 4,184 DC50 (potency) measurements and 2,377 Dmax (maximal degradation efficacy) measurements, covering 3,514 unique PROTACs, 164 POI (protein of interest) targets, 9 E3 ligases, and 155 cell lines. It offers four configurations: default (all endpoints combined), DC50 (potency measurements only), Dmax (efficacy measurements only), and multitask (paired DC50 and Dmax data for the same PROTAC/assay, used for binary activity classification). Each configuration features comprehensive columns such as SMILES (canonical representation of PROTACs), POI and ligase sequences with UniProt annotations, cell line information (standardized via Cellosaurus), measurement values (including units, errors, ranges), experimental conditions (e.g., assay time, concentration), source database references, and structural clustering identifiers for cross-validation and hold-out sets. The dataset includes a structurally dissimilar hold-out set (~10% of data) for rigorous benchmarking and supports regression (pDC50, Dmax) and binary classification tasks. Benchmark results show that pDC50 is more predictable (R² of 0.66) than Dmax (R² of 0.36) on the hold-out set, and classical tree-based methods (e.g., XGBoost) outperform domain-specific GNNs in classification. TACK aims to advance ML research for PROTAC design, providing pre-trained models and a statistical evaluation framework.



