Diffusion models for binder design
The modern de novo binder recipe has three steps: generate a backbone against a chosen patch on the target, design a sequence for that backbone, then predict whether the sequence folds back into it. RFdiffusion, ProteinMPNN and a structure predictor are the canonical trio, and BindCraft and similar pipelines package the loop.
Generation
Diffusion starts from noise and denoises toward a protein-like backbone, conditioned on the target structure and on the residues you nominate as the hotspot. The hotspot choice is the single most consequential input: the method does not find a good epitope, it builds toward the patch you gave it. Pick that patch from biology, from a known interaction interface, or from a pocket you have reason to believe matters.
Scaffold length and topology are the other levers. Short helical bundles are what these methods do best, which is why so many published de novo binders are three-helix bundles. Loop-heavy interfaces and binders that must reach into a narrow cleft remain much harder, and that is exactly where antibody and VHH formats keep their advantage.
Sequence design
ProteinMPNN takes the generated backbone and writes sequences likely to fold into it, sampling at a temperature you choose. Low temperature gives conservative, high-confidence sequences; higher temperature gives diversity at the cost of foldability. Designing several sequences per backbone and filtering afterward is cheaper than trying to get one right.
Fixing interface residues during design, so the contacts the backbone was built around survive, is standard practice and it matters.
The filter cascade
Nothing in generation guarantees a design folds, binds or behaves. The filters that carry weight, roughly in the order they are worth running:
- Self-consistency. Predict the structure of the designed sequence alone and compare it to the backbone it was designed for. Designs whose prediction does not recover the intended fold are discarded, and this removes a large fraction.
- Complex prediction. Co-fold the design with the target and look at interface confidence and whether the predicted pose is the intended one, not just any pose.
- Interface quality. Buried surface area, shape complementarity, unsatisfied buried polar atoms, and the hydrophobic content of the interface.
- Developability. Free cysteines, aggregation-prone surface, extreme charge. Cheap to check and expensive to ignore.
- Simulation on the shortlist, to see whether the pose survives water and temperature.
What to expect experimentally
Published campaigns that follow this recipe typically order tens to hundreds of designs and find that a small percentage bind at useful affinity, with the rest not binding at all. That is a good outcome by historical standards and it is still a screening exercise, not a design-and-ship pipeline. Plan the build and test capacity accordingly, and treat any computational score as a way to spend fewer wells rather than as a prediction of success.
The failure mode to watch for is a pipeline tuned to produce designs that pass its own filters. Self-consistency and interface confidence are optimization targets that can be gamed by the generator, and a panel with perfect scores and no binders is a common result of over-filtering on metrics nobody validated against wet data.