Microscopic pKa Prediction

Python API · stjames models · API example

How it works

Predict the pKa for adding or removing a proton at individual sites in the supplied molecular state. Results identify conjugate acids and bases, their sites, and the strongest acid and base found. A protonation result reports the pKa of the resulting conjugate acid.

Settings

  • Method: the web default is "Rowan pKa (Wagen 2026) - g-xTB-Based Aqueous pKa (All Elements)." The other options are "Rowan pKa (Wagen 2024) - NNP-Based Aqueous pKa," "Starling - ML-Based Aqueous pKa," and "Chemprop (Nevolianis 2025) - ML-Based Aqueous and Nonaqueous pKa."
  • Solvent: water for g-xTB, AIMNet2, and Starling. Chemprop also supports DMSO, DMF, acetonitrile, methanol, ethanol, ethylene glycol, and N-methylpyrrolidone.
  • Min pKa and Max pKa: default to 2 and 12. These limit displayed values; the g-xTB and AIMNet2 methods also screen unlikely sites to save time.
  • Protonate elements and Deprotonate elements: default to N and N, O, S, respectively. Chemprop supports deprotonation only.
  • Protonate indices and Deprotonate indices: add specific sites using atom numbers starting at 1, including ranges such as 7-9. These supplement the element lists; clear those lists to select only numbered sites. Indices are disabled for Starling and Chemprop.
  • Mode: controls conformer sampling for g-xTB and AIMNet2; disabled for Starling and Chemprop.

Modes

"Careful" is the web default and samples most broadly. "Rapid" and "Reckless" reduce cost by retaining fewer conformers. Both g-xTB and AIMNet2 use these settings:

ModeCarefulRapidReckless
screening buffer (pKa units)1055
openconf Monte Carlo steps400200100
maximum conformers retained1031
energy window (kcal/mol)101010

Notes

For whole-molecule ionization and microstate populations, use macroscopic pKa prediction, especially when several sites have similar pKa values.

Starling and Chemprop predict directly from molecular connectivity. The g-xTB and AIMNet2 methods instead calculate conformer free energies and require closed-shell, connected molecular structures. g-xTB supports a wider range of elements, including metals; AIMNet2 is limited to its supported elements.

Protonation, deprotonation, and conformer generation

An overview of how Rowan predicts pKa values

An overview of how Rowan predicts pKa values from Wagen 2024.

The conformer methods add or remove one proton at eligible sites and screen rough pKa estimates before detailed calculations. g-xTB reuses parent conformers for conjugate states; AIMNet2 searches each state independently. Both optimize without implicit solvent, then add GFN2-xTB/CPCM-X water solvation corrections. g-xTB uses one representative GFN2 vibrational correction per state; AIMNet2 evaluates vibrations for each retained conformer.

ΔGs to pKas

Conformer contributions are combined into ensemble free energies at 298.15 K, following the approach of Pracht and Grimme. Empirical calibration converts the acid–base free-energy difference to pKa: functional-group-specific linear relationships for g-xTB, and a quadratic relationship with element/valence corrections for AIMNet2.

Submission video

Benchmarks and validation

Accuracy

On a variety of benchmarks (also described in the preprint), Rowan pKa displays mean absolute errors of around 1 pKa unit, and decent predictive accuracy as assessed by Kendall's τ and R2 values. Here's the AIMNet2 method's performance on the SAMPL7 benchmark set:

Rowan pKa performance on SAMPL7

Rowan pKa prediction workflow's performance on SAMPL7

For a full list of all assays surveyed, see the preprint.

We reinvestigated a variety of benchmark systems with the new g-xTB method after development was finished (without using the benchmarks to tune any scaling factors), and found improved accuracy across the board. On the Rombouts set of tricyclic amidine BACE1 inhibitors, for instance, g-xTB shrinks the MAE from 1.06 to 0.59 pKa units.

Rowan pKa performance on the Rombouts set of tricyclic amidine BACE1
inhibitors

Rowan pKa prediction workflow's performance on the Rombouts set of tricyclic amidine BACE1 inhibitors

We also see improvements on the Miller–Doukas–Seydel folate-inhibitor dataset:

Rowan pKa performance on the Miller–Doukas–Seydel folate-inhibitor dataset
inhibitors

Rowan pKa prediction workflow's performance on the Miller–Doukas–Seydel folate-inhibitor dataset

And on the Müller dataset examining the effect of α- and β-substituents (e.g. oxetanes) on cyclic amine basicity:

Rowan pKa performance on Müller dataset

Rowan pKa prediction workflow's performance on the Müller cyclic amine basicity dataset

Further reading