LogP Prediction

Python API · stjames models · API example

How it works

This workflow predicts logP, the base-10 logarithm of a neutral compound's equilibrium concentration ratio between 1-octanol and water. Use it to compare lipophilicity during molecular design or screening. Each result is a single, dimensionless logP value.

Settings

The "Method" selector offers three choices:

  • Chemprop (SangsterLogP) - ML logP: The default. A machine-learning model trained on curated experimental data from the SangsterLogP dataset. It predicts from molecular connectivity and is a practical starting point for fast screening. Reliability depends on how well the training data represent your chemistry.
  • Crippen - Rule-Based logP: RDKit's Wildman–Crippen atom-contribution model. It gives a fast, reproducible baseline from molecular connectivity, without explicitly modeling conformations or solvent thermodynamics.
  • COSMO-RS - Physics-Based logP: Samples molecular conformations and uses quantum-chemical surfaces to estimate transfer between water and wet octanol at 298.15 K (25 °C). It combines conformer contributions according to their equilibrium populations in water. Choose it when conformational and solvent effects warrant the additional calculation time; it is substantially slower than the other methods.

Temperature and solvent composition are fixed for COSMO-RS; the form exposes only the method choice.

Notes

Larger values indicate greater preference for octanol. For example, logP = 2 corresponds to a 100-fold higher concentration in octanol than in water. Negative values indicate a preference for water.

Use the neutral molecular structure rather than a salt or disconnected mixture. LogP differs from logD, which includes ionized species at a specified pH; this workflow does not predict a pH-dependent distribution curve or enumerate tautomers. Treat results as predictions for the submitted structure. COSMO-RS provides a different physical approach, but using it outside the learned model's domain does not guarantee better accuracy.

Benchmarks and validation

Predicted versus experimental logP on the SangsterLogP prospective set: Chemprop R squared 0.765 and RDKit Crippen R squared 0.577.

Performance comparison of Chemprop D-MPNN and RDKit Crippen on the prospective set from the SangsterLogP paper. Figures from our logP update.

COSMO-RS predicted versus experimental logP for approved drugs from ChEMBL, with R squared 0.676.

LogP prediction with COSMO-RS on a set of approved drugs derived from the ChEMBL dataset.

Further reading