[Paper Review] Assessing interaction recovery of predicted protein-ligand poses
This paper introduces protein-ligand interaction fingerprint (PLIF) recovery as a critical metric for evaluating protein-ligand pose prediction models, demonstrating that despite low RMSD and PoseBuster validity, many machine learning-based docking and cofolding models fail to recapitulate key interactions like hydrogen and halogen bonds. The study reveals that classical docking methods significantly outperform ML approaches in interaction recovery, highlighting the need for explicit pharmacophore-aware loss functions in model training.
The field of protein-ligand pose prediction has seen significant advances in recent years, with machine learning-based methods now being commonly used in lieu of classical docking methods or even to predict all-atom protein-ligand complex structures. Most contemporary studies focus on the accuracy and physical plausibility of ligand placement to determine pose quality, often neglecting a direct assessment of the interactions observed with the protein. In this work, we demonstrate that ignoring protein-ligand interaction fingerprints can lead to overestimation of model performance, most notably in recent protein-ligand cofolding models which often fail to recapitulate key interactions.
Motivation & Objective
- To address the overreliance on RMSD and PoseBuster validity as sole metrics for pose prediction quality.
- To demonstrate that interaction recovery—particularly of key interactions like hydrogen and halogen bonds—is a necessary but often overlooked criterion for pose validity.
- To evaluate the performance of classical docking, ML docking, and cofolding models in recovering ground truth protein-ligand interaction fingerprints (PLIFs).
- To advocate for the integration of explicit pharmacophore or PLIF-sensitive loss functions in machine learning models to improve biological relevance of predicted poses.
Proposed method
- Protein-ligand interaction fingerprints (PLIFs) were calculated using the ProLIF package (v2.0.3), focusing on specific interaction types: hydrogen bonds, halogen bonds, π-stacking, cation-π, π-cation, and ionic interactions.
- Custom distance thresholds were applied: 3.7 Å for hydrogen bonds, 5.5 Å for cation-π, and 5.0 Å for ionic interactions, with all other parameters set to default.
- PLIF recovery was measured as the ratio of correctly recovered interactions to ground truth interactions in the crystal structure, with recall calculated per interaction type.
- Models were evaluated on the PoseBusters test set (308 complexes), with RMSD and PoseBuster validity used as baseline metrics, while PLIF recovery served as the primary novel metric.
- The analysis compared classical docking (GOLD), ML docking (DiffDock-L), and cofolding models (RoseTTAFold-AllAtom, Umol, etc.) across interaction recovery and geometric accuracy.
- A modified definition of PoseBuster validity was used, excluding ligand RMSD to isolate the contribution of interaction recovery to pose quality.

Experimental results
Research questions
- RQ1To what extent do state-of-the-art machine learning-based protein-ligand pose prediction models recover key protein-ligand interactions compared to crystal structures?
- RQ2How does PLIF recovery correlate with RMSD and PoseBuster validity, and can it reveal performance limitations missed by these standard metrics?
- RQ3Why do cofolding models, despite achieving low RMSD, fail to recover biologically relevant interactions like hydrogen bonds and halogen bonds?
- RQ4Can the integration of explicit pharmacophore or interaction-aware loss functions improve the biological relevance of ML-predicted poses?
Key findings
- Classical docking (GOLD) recovered 100% of hydrogen bonds and all key interactions in the 6M2B complex, while DiffDock-L missed the halogen bond and altered the hydrogen bonding pattern.
- DiffDock-L recovered 75% of the PLIFs in the 6M2B complex, indicating partial but incomplete recapitulation of key interactions despite a low RMSD of 1.5 Å.
- RoseTTAFold-AllAtom achieved an RMSD of 2.19 Å and failed to recover any of the ground truth interactions, despite being PoseBuster-valid.
- On average, ML docking and cofolding models recovered significantly fewer hydrogen bonds and π-stacking interactions than classical docking, with the latter consistently outperforming ML methods in interaction recovery.
- The study found a weak correlation between PLIF recovery and RMSD, indicating that low RMSD does not guarantee recovery of biologically relevant interactions.
- Cofolding models often produce poses with steric clashes and incorrect orientations, leading to spurious cationic interactions that replace genuine cation-π or π-stacking interactions found in the crystal structure.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.