tiny226/SAVI-2020-NOVA
收藏资源简介:
SAVI 2020数据集是一个全面的计算生成的有机化合物集合,包含超过15亿个旨在易于实际合成的分子。该数据集通过基于有机化学知识的专家系统类型规则创建,代表了药物发现和化学研究中最大的公开可用的合成可及虚拟化合物数据库之一。数据库使用在CHMTRN/PATRAN编程语言中编码的53个精心挑选的化学转化规则生成,这些规则最初为LHASA(逻辑和启发式应用于合成分析)逆合成分析系统开发。这些转化应用于Enamine的大约152,532个商业可用构建块,通过单步、双反应物合成路径创建分子。SAVI化合物对于虚拟筛选、药物发现和计算化学应用特别有价值,因为每个分子都附带:提议的单步合成路线、预测的合成可及性评分、全面的分子属性注释以及基于药物相似性标准的质量评估。
The Synthetically Accessible Virtual Inventory (SAVI) 2020 dataset is a comprehensive collection of over 1.5 billion computationally generated organic compounds designed to be easily and practically synthesizable. Created through expert-system type rules derived from established organic chemistry knowledge, SAVI represents one of the largest publicly available databases of synthetically accessible virtual compounds for drug discovery and chemical research. The database was generated using 53 carefully selected chemical transformation rules encoded in the CHMTRN/PATRAN programming languages, originally developed for the LHASA (Logic and Heuristics Applied to Synthetic Analysis) retrosynthetic analysis system. These transforms were applied to approximately 152,532 commercially available building blocks from Enamine to create molecules through single-step, two-reactant synthesis pathways. SAVI compounds are particularly valuable for virtual screening, drug discovery, and computational chemistry applications because each molecule comes with: A proposed single-step synthetic route, Predicted synthetic accessibility scores, Comprehensive molecular property annotations, Quality assessments based on drug-likeness criteria.



