Protein–Ligand Co-Folding

Python API · stjames models · API example · Constraints example

How it works

Predict the three-dimensional structure of a protein–ligand complex from protein sequences and ligand chemistry, without supplying a pre-docked pose. Rowan returns predicted complexes and model confidence scores; supported models also predict binding affinity. Optional refinement and ligand validation help assess the predicted binding pose.

Settings

  • Inputs: supply protein sequences, ligands, and cofactors. DNA and RNA are supported by selected models; Rowan's Boltz-1 and DeCAF-Boltz workflows do not accept them.
  • Model: choose "Chai-1r," "Boltz-1," "Boltz-2," "Boltz-2.1," "OpenFold3-preview," or "DeCAF-Boltz." The web form starts with Boltz-2.
  • Samples: "Num samples" controls how many candidate structures are predicted. Compare their confidence and ligand geometry.
  • Alignments: "Use MSA server?" requests multiple sequence alignments (MSAs) and is enabled initially. Alignments provide evolutionary information to support structure prediction.
  • Guidance: "Use potentials?" enables physical guidance in supported Boltz models and starts off. Contact and pocket constraints are available for Boltz-2 and Boltz-2.1; templates are also available for OpenFold3-preview. Constraints guide predictions and should be checked in the resulting structure.
  • Affinity: "Predict binding affinity?" is available for Boltz-2 and Boltz-2.1 with ligands. Their reported affinity metrics differ; compare results within the same model.
  • Refinement: "Local opt. and MM/GBSA?" locally relaxes the primary ligand pose and estimates its interaction energy. The web form requires affinity prediction to be enabled first.
  • Strain: "Compute strain?" adds a conformer search to estimate the ligand's energetic penalty for adopting the predicted geometry. It requires local optimization and adds runtime.

Notes

Modified polymer residues can be specified with a sequence position and a PDB Chemical Component Dictionary (CCD) code, such as SEP for phosphoserine or TPO for phosphothreonine, where supported by the model.

For the designated primary ligand, Rowan checks extracted poses against the requested chemistry and stereochemistry and applies PoseBusters geometry and clash checks. These checks can flag implausible poses; passing them or obtaining high model confidence does not establish experimental binding.

Refinement applies MM/GBSA to the top-ranked sample. Optional strain calculations then use restrained GFN2-xTB/ALPB optimization and g-xTB/CPCM-X energies in water, relative to the lowest-energy conformer found. Strain and MM/GBSA scores are approximate energy estimates, distinct from model-predicted affinity. Without refinement, extracted poses are checked without an energy calculation. Unsupported protein or cofactor chemistry can prevent MM/GBSA scoring; covalent bond constraints disable MM/GBSA and strain processing.

Boltz workflows warn when ligands exceed 50 heavy atoms because large ligands are underrepresented in training. Inspect warnings and predicted poses before using them for downstream work.

Submission video

Benchmarks and validation

Runs N’ Poses benchmark comparing protein–ligand pose prediction success across AlphaFold3, OpenFold3-preview-2, Protenix, Boltz-1, Boltz-2, and Chai-1, grouped by similarity to training data.

Figure 2 from the OpenFold3-preview-2 technical report. Bins are ordered by similarity to the training data, from lowest (left) to highest (right).

Further reading