Better Geometries, Better Electronics
Fine-tuning AIMNet2 with NVIDIA ALCHEMI to power covalent drug design
How Expedition Medicines used the NVIDIA ALCHEMI Toolkit to fine-tune a neural network potential and relax millions of reactive structures to DFT quality — 15 million conformers per $1,000 of compute.
Drugging the undruggable, one covalent bond at a time
Expedition Medicines designs covalent small molecules against targets that have resisted conventional drug discovery. Where reversible binders fail, a precisely placed covalent bond can lock a compound onto a target and modulate its function. The hard part is precision: the electrophile must react with the intended residue (often a target cysteine) while ignoring the tens of thousands of competing nucleophiles in the cell. That selectivity is governed at the ångström scale by the local electronics of the electrophile, for example its frontier orbital energies, partial charges, and electrostatic potential at the reactive center. This is why our models generate compounds informed by their electronics, not just their atoms: new electronic profiles unlock new targets.
Why Expedition Medicines’ generative chemistry models need DFT
Expedition trains generative foundation models, including latent diffusion models, that propose new electrophiles conditioned on target sites and off-target reactivity. We train them on more than 1.4 billion measured covalent bond-formation events spanning the proteome. We cotrain on large-scale synthetic quantum-chemistry data: frontier orbital energies, atomic charges, and electrostatic descriptors computed at a hybrid density-functional level of theory. That second, synthetic source grounds the models in the physics of how electrons rearrange when a bond forms, so they generalize beyond the reactions we have measured directly. But DFT electronic properties are exquisitely sensitive to nuclear coordinates. On a poorly relaxed structure, for example a strained ring, a misplaced rotamer, the descriptors are still accurate, but they describe an artificial, high-energy state instead of reality. Therefore, every structure in the co-training set must first be relaxed to a faithful energy minimum.
Geometry optimization is the bottleneck
Relaxing geometries with DFT is the rate-limiting, cost-dominating step of building our co-training corpus. AIMNet2 removes it: a neural network potential trained on roughly 107 hybrid-DFT calculations spanning up to 14 elements in neutral and charged states. For geometry optimization, it matches DFT within ~1 – 2 kcal/mol on conformational energies while running about five orders of magnitude faster.
With one catch. AIMNet2 was trained on broadly representative organic chemistry. Our electrophiles are not broadly representative; they can occupy reactive, charge-polarized regions of chemical space that are under-represented in AIMNet2’s original training set — next generation reactive chemistries. On these molecules, vanilla AIMNet2 drifts, relaxing to geometries dramatically higher in energy, and the resulting electronic descriptors carry that error forward into our models. So, we fine-tuned AIMNet2.
Figure 1. Geometry optimization with a fine-tuned AIMNet2 (AIMNet2-FT) sits between chemistry generation and single-point DFT, feeding the electronic co-training set for our reactivity models.
Fine-tuning AIMNet2 with the NVIDIA ALCHEMI Toolkit
Using the NVIDIA ALCHEMI Toolkit, we fine-tuned AIMNet2 using more than 13,000 compounds representing a diverse subset of our reactive chemistries, with reference energies and forces computed at wB97X-d3/def2-TZVPP level of theory. The goal was to retain AIMNet2’s speed while bringing its energy and force predictions into close agreement with DFT in the regions of chemical space most relevant to our compounds. Fine-tuning ran on NVIDIA GPUs using the ALCHEMI Toolkit, which provided a streamlined, configuration-driven workflow for setting up, validating, and scaling training runs without extensive custom infrastructure. At inference we maintained greater than 90% GPU utilization across a range of GPU architectures owing to ALCHEMI Toolkit’s batching systems.
Results
The fine-tuned model finds lower-energy, more DFT-faithful geometries than vanilla AIMNet2 on our reactive compounds, and it does so at the same speed.
Fine-tuning improved agreement with DFT on our reactive-first chemistry without degrading performance. On conventional covalent scaffolds the fine-tuned and base models are equivalent, with median |ΔE| of 1.829 and 0.993 kcal/mol across two representative scaffolds against 1.841 and 1.093 for vanilla AIMNet2; on the next generation reactive chemistries, the median falls from 3.185 to 1.997 kcal/mol.
Importantly, the tails separate far more sharply than the medians (Figure 1). On our chemistry the base model’s 90th-percentile |ΔE| exceeds 300 kcal/mol, whereas the fine-tuned model’s p90 is below 4 kcal/mol. GFN2-xTB is less accurate at the median on both sets (2.751 / 2.712 and 5.926 kcal/mol) but lacks a long tail, consistent with bounded error from explicit electronic structure rather than unbounded extrapolation error. Optimization convergence follows the same pattern: 89.6% for vanilla AIMNet2 against 100% for both the fine-tuned model and GFN2-xTB.
Table 1. Benchmark of fine-tuned vs. vanilla AIMNet2 vs GFN2-xTB against reference DFT *Calculations exclude molecules that deviate more than 20kcal/mol in energy from DFT.
AIMNet2 unlocked accuracy at scale
At roughly 100 ms per molecule on a GPU, the fine-tuned AIMNet2 model makes high-throughput molecular optimization practical at a fundamentally different scale. It is approximately two orders of magnitude faster than GFN2-xTB on CPU and produces about 15 million geometry-optimized conformers per $1,000 of compute, compared with roughly 900,000 for GFN2-xTB. This represents a 17-fold cost advantage, while also providing higher accuracy on our chemistry. We have already optimized millions of molecules using this approach. Because geometries reliably converge to high-quality minima on the first pass, the need for subsequent DFT-level refinement has largely been eliminated, removing a major computational bottleneck and further increasing the throughput of each successive co-training cycle.
Figure 2. Median and 90th-percentile |ΔE| against DFT for conventional covalent chemistry (top, linear axis) and Expedition’s reaction-first chemistry (bottom, log axis). On our chemistry only the fine-tuned model keeps both statistics tight; vanilla AIMNet2’s tail extends nearly two orders of magnitude beyond its median.
“The next generation of drug discovery will be powered by AI models that learn from both molecular structure and the electronic properties that govern reactivity. By fine-tuning AIMNet2 with the NVIDIA ALCHEMI Toolkit, Expedition Medicines can generate DFT-quality geometries across millions of compounds, creating scalable, high-fidelity training data for covalent drug design.
—Geetika Gupta, Sr. Director, AI for Science, NVIDIA
Why this matters
Higher-quality geometries mean higher-quality electronic descriptors, which mean a cleaner learning signal for our generative models. Faster, cheaper optimization means we can refresh and expand the pre-training set as our chemistry evolves, rather than treat it as a one-time, budget-limited effort. This lets us scale quantum chemistry calculations to tens or hundreds of millions of conformers and generate DFT-derived synthetic data for millions of molecules to train our next generation of generative models. The combination tightens the loop between what our models propose and what is electronically and synthetically real: the loop that ultimately produces selective covalent drugs against targets others have written off.
Acknowledgements. This work was carried out by Expedition Medicines in collaboration with NVIDIA. AIMNet2 was developed by Anstine, Zubatyuk, and Isayev.
Reference: Anstine, Zubatyuk & Isayev, “AIMNet2: a neural network potential…,” Chemical Science (2025).