ByteDance-Seed/SCFbench
收藏资源简介:
--- license: apache-2.0 --- <div align="center"> 👋 Hi, everyone! <br> We are <b>ByteDance Seed team.</b> </div> <p align="center"> You can get to know us better through the following channels👇 <br> <a href="https://seed.bytedance.com/"> <img src="https://img.shields.io/badge/Website-%231e37ff?style=for-the-badge&logo=bytedance&logoColor=white"></a> <a href="https://github.com/user-attachments/assets/5793e67c-79bb-4a59-811a-fcc7ed510bd4"> <img src="https://img.shields.io/badge/WeChat-07C160?style=for-the-badge&logo=wechat&logoColor=white"></a> <a href="https://www.xiaohongshu.com/user/profile/668e7e15000000000303157d?xsec_token=ABl2-aqekpytY6A8TuxjrwnZskU-6BsMRE_ufQQaSAvjc%3D&xsec_source=pc_search"> <img src="https://img.shields.io/badge/Xiaohongshu-%23FF2442?style=for-the-badge&logo=xiaohongshu&logoColor=white"></a> <a href="https://www.zhihu.com/org/dou-bao-da-mo-xing-tuan-dui/"> <img src="https://img.shields.io/badge/zhihu-%230084FF?style=for-the-badge&logo=zhihu&logoColor=white"></a> </p>  # Towards A Universally Transferable Acceleration Method for Density Functional Theory Zhe Liu, Yuyan Ni, Zhichen Pu, Qiming Sun, Siyuan Liu & Wen Yan https://arxiv.org/abs/2509.25724 # TL;DR We propose a framework for accelerating DFT calculations. We train E(3)-equivariant neural networks to predict the expansion coefficients of the electron density in an auxiliary basis, and use the prediction to construct an initial guess for the SCF process. This approach exhibits superior transferability in various aspects. # Changelog ## 2026.3.7 * The full evaluation code is released. * The train/valid/test split of the main dataset is released. * The model weights of the NequIP model are released. ## 2025.12.1 * First release. # Contents The repo currently contains the following contents: * The full SCFbench dataset. * The data pipeline for the SCFbench dataset. * The PyTorch `nn.Module` of the species-wise linear layer for the prediction of the electron density coefficients. * The NequIP model code and weights with the species-wise linear layer. * Example code for computing the density coefficients from a density matrix. * The full evaluation code. # Requirements * torch * e3nn * pyscf * lmdb * numpy>1.26 * nequip (if you want to use the NequIP model) # Dataset Usage The `dataset` folder of the repo contains the `main` dataset (the dataset for training, validation and in-distribution testing) and the `ood-test` dataset. Each dataset contains several `parts`, each of which corresponds to a specific piece of information. The parts are: * `base`: the basic information of the molecule, including atomic numbers, coordinates, etc. * `dm`: the density matrix of the molecule. * `fock`: the Hamiltonian (fock) matrix of the molecule. * `auxdensity.denfit`: the density coefficients on def2-universal-jfit. * `auxdensity.denfit.etb2.0`: the density coefficients on the ETB basis of def2-svp with $\beta=2.0$. * `auxdensity.denfit.etb1.5`: the density coefficients on the ETB basis of def2-svp with $\beta=1.5$. Example usage: ```python from dataset import SCFBenchDataset # Loading base info (atomic numbers, coordinates, etc.), density matrix, Hamiltonian (fock) matrix and the density coefficients on def2-universal-jfit. parts_to_load = ['base', 'dm', 'fock', 'auxdensity.denfit'] dataset = SCFBenchDataset(data_root='dataset/main', parts_to_load=parts_to_load) dataset[0].keys() # Loading the base info and the density coefficients on the ETB basis of def2-svp with $\beta=1.5$. parts_to_load = ['base', 'auxdensity.denfit.etb1.5'] dataset = SCFBenchDataset(data_root='dataset/ood-test', parts_to_load=parts_to_load, auxbasis='etb:def2-svp:1.5') dataset[0].keys() # for the raw data, use the underlying dataset dataset.dataset[0].keys() ``` # Evaluation The full evaluation code is provided in `evaluate_scf_gpu.py`. To evaluate the provided NequIP model checkpoint `nequip_L_jfit.ckpt`, run: ```bash # evaluate on the test split of the main dataset (ID setting); note that you can also specify --num-shards and --shard-index to only run the evaluation on a subset python evaluate_scf_gpu.py --ckpt nequip_L_jfit.ckpt --data-root dataset/main --split test --output id_test.csv # evaluate on the ood-test dataset python evaluate_scf_gpu.py --ckpt nequip_L_jfit.ckpt --data-root dataset/ood-test --split no --output ood_test.csv # evaluate on the ood-test dataset, with a XC/basis transfer setting python evaluate_scf_gpu.py --ckpt nequip_L_jfit.ckpt --data-root dataset/ood-test --split no --xc blyp --transfer-basis def2-tzvp --output ood_test_transfer.csv ``` # Known Issues * The species-wise linear readout layer may cause unstable model forward time on GPU in some software environments. We recommend implementing the prediction using a uniform padded basis for all species, just like how Hamiltonian prediction models (e.g., QHNet) predict the Hamiltonian matrix. We have tested this method internally but would like to keep the code in this repo consistent with our paper for reproducibility. # Citing SCFbench If you use SCFbench in your research, please cite: ```latex @misc{liu2025universallytransferableaccelerationmethod, title={Towards A Universally Transferable Acceleration Method for Density Functional Theory}, author={Zhe Liu and Yuyan Ni and Zhichen Pu and Qiming Sun and Siyuan Liu and Wen Yan}, year={2025}, eprint={2509.25724}, archivePrefix={arXiv}, primaryClass={physics.chem-ph}, url={https://arxiv.org/abs/2509.25724}, } ``` ## License Models are licensed under the [Apache License, Version 2.0](http://www.apache.org/licenses/LICENSE-2.0). The dataset is a derivative of [ChEMBL](https://www.ebi.ac.uk/chembl/), used under [CC BY-SA 3.0](https://creativecommons.org/licenses/by-sa/3.0/). Our modified version, the SCFbench dataset, is also licensed under [CC BY-SA 3.0](https://creativecommons.org/licenses/by-sa/3.0/). ## About [ByteDance Seed Team](https://seed.bytedance.com/) Founded in 2023, ByteDance Seed Team is dedicated to crafting the industry's most advanced AI foundation models. The team aspires to become a world-class research team and make significant contributions to the advancement of science and society.
--- license: apache-2.0 --- <div align="center"> 👋 大家好! <br> 我们是<b>字节跳动Seed团队(ByteDance Seed team)</b> </div> <p align="center"> 您可通过以下渠道进一步了解我们👇 <br> <a href="https://seed.bytedance.com/"> <img src="https://img.shields.io/badge/Website-%231e37ff?style=for-the-badge&logo=bytedance&logoColor=white"></a> <a href="https://github.com/user-attachments/assets/5793e67c-79bb-4a59-811a-fcc7ed510bd4"> <img src="https://img.shields.io/badge/WeChat-07C160?style=for-the-badge&logo=wechat&logoColor=white"></a> <a href="https://www.xiaohongshu.com/user/profile/668e7e15000000000303157d?xsec_token=ABl2-aqekpytY6A8TuxjrwnZskU-6BsMRE_ufQQaSAvjc%3D&xsec_source=pc_search"> <img src="https://img.shields.io/badge/Xiaohongshu-%23FF2442?style=for-the-badge&logo=xiaohongshu&logoColor=white"></a> <a href="https://www.zhihu.com/org/dou-bao-da-mo-xing-tuan-dui/"> <img src="https://img.shields.io/badge/zhihu-%230084FF?style=for-the-badge&logo=zhihu&logoColor=white"></a> </p>  # 面向密度泛函理论的通用可迁移加速方法 Zhe Liu, Yuyan Ni, Zhichen Pu, Qiming Sun, Siyuan Liu & Wen Yan https://arxiv.org/abs/2509.25724 # 核心要点(TL;DR) 我们提出了一种用于加速密度泛函理论(Density Functional Theory, DFT)计算的框架。我们训练E(3)等变神经网络来预测辅助基组下电子密度的展开系数,并利用该预测结果为自洽场(Self-Consistent Field, SCF)过程构建初始猜测。该方法在多个维度上展现出优异的可迁移性。 # 更新日志(Changelog) ## 2026.3.7 * 完整的评估代码正式发布 * 主数据集的训练/验证/测试划分集已发布 * NequIP模型的权重文件已发布 ## 2025.12.1 * 首次发布 # 仓库内容 本代码仓库当前包含以下内容: * 完整的SCFbench数据集 * SCFbench数据集的数据处理流水线 * 用于预测电子密度系数的物种线性层的PyTorch `nn.Module` * 搭载物种线性层的NequIP模型代码与权重 * 从密度矩阵计算密度系数的示例代码 * 完整的评估代码 # 依赖要求 * torch * e3nn * pyscf * lmdb * numpy>1.26 * nequip(若需使用NequIP模型) # 数据集使用方法 仓库的`dataset`文件夹包含`main`数据集(用于训练、验证和分布内测试)与`ood-test`数据集。 每个数据集包含若干`parts`,每个`part`对应一类特定信息,具体如下: * `base`:分子的基础信息,包括原子序数、坐标等 * `dm`:分子的密度矩阵 * `fock`:分子的哈密顿(Fock)矩阵 * `auxdensity.denfit`:def2-universal-jfit基组下的密度系数 * `auxdensity.denfit.etb2.0`:def2-svp基组的ETB基下($eta=2.0$)的密度系数 * `auxdensity.denfit.etb1.5`:def2-svp基组的ETB基下($eta=1.5$)的密度系数 示例使用代码: python from dataset import SCFBenchDataset # 加载基础信息(原子序数、坐标等)、密度矩阵、哈密顿(Fock)矩阵以及def2-universal-jfit基组下的密度系数 parts_to_load = ['base', 'dm', 'fock', 'auxdensity.denfit'] dataset = SCFBenchDataset(data_root='dataset/main', parts_to_load=parts_to_load) dataset[0].keys() # 加载基础信息以及def2-svp基组ETB基下($eta=1.5$)的密度系数 parts_to_load = ['base', 'auxdensity.denfit.etb1.5'] dataset = SCFBenchDataset(data_root='dataset/ood-test', parts_to_load=parts_to_load, auxbasis='etb:def2-svp:1.5') dataset[0].keys() # 访问原始数据,请使用底层dataset对象 dataset.dataset[0].keys() # 模型评估 完整的评估代码位于`evaluate_scf_gpu.py`中。若需对提供的NequIP模型检查点`nequip_L_jfit.ckpt`进行评估,可运行以下命令: bash # 在主数据集的测试划分(分布内设置)上进行评估;注:也可通过指定--num-shards和--shard-index仅对部分数据运行评估 python evaluate_scf_gpu.py --ckpt nequip_L_jfit.ckpt --data-root dataset/main --split test --output id_test.csv # 在ood-test数据集上进行评估 python evaluate_scf_gpu.py --ckpt nequip_L_jfit.ckpt --data-root dataset/ood-test --split no --output ood_test.csv # 在ood-test数据集上进行交换泛函/基组的迁移设置下的评估 python evaluate_scf_gpu.py --ckpt nequip_L_jfit.ckpt --data-root dataset/ood-test --split no --xc blyp --transfer-basis def2-tzvp --output ood_test_transfer.csv # 已知问题 * 物种线性读出层在部分软件环境下可能会导致GPU上的模型前向时间不稳定。我们建议使用针对所有物种的统一填充基组来实现预测,正如哈密顿预测模型(如QHNet)预测哈密顿矩阵的方式。我们已在内部测试过该方法,但为保证可复现性,将保留本仓库代码与论文中的一致。 # 引用SCFbench 若在研究中使用SCFbench,请引用如下文献: latex @misc{liu2025universallytransferableaccelerationmethod, title={Towards A Universally Transferable Acceleration Method for Density Functional Theory}, author={Zhe Liu and Yuyan Ni and Zhichen Pu and Qiming Sun and Siyuan Liu and Wen Yan}, year={2025}, eprint={2509.25724}, archivePrefix={arXiv}, primaryClass={physics.chem-ph}, url={https://arxiv.org/abs/2509.25724}, } ## 许可证 模型采用[Apache许可证2.0版](http://www.apache.org/licenses/LICENSE-2.0)进行授权。 本数据集衍生自[ChEMBL](https://www.ebi.ac.uk/chembl/),遵循[CC BY-SA 3.0](https://creativecommons.org/licenses/by-sa/3.0/)协议进行使用。 我们修改后的SCFbench数据集同样采用[CC BY-SA 3.0](https://creativecommons.org/licenses/by-sa/3.0/)协议进行授权。 ## 关于字节跳动Seed团队(ByteDance Seed Team) 字节跳动Seed团队成立于2023年,致力于打造行业领先的人工智能基础模型。团队立志成为世界级研究团队,为科学与社会的进步做出重要贡献。



