Co-folding and its limits
Co-folding is structure prediction run on more than one chain at once: you give the model a binder and a target and it returns a complex. The models in common use, including AlphaFold-Multimer, AlphaFold3 and the open implementations that followed, have made this routine enough that it now sits in most design pipelines as a filter.
It is worth being precise about what it delivers, because the failure mode is expensive: a confident, wrong complex looks exactly like a confident, right one.
What the confidence scores mean
Two numbers do most of the work.
PAE, the predicted aligned error, is a matrix. The part that matters for a complex is the off-diagonal block, which estimates the error in one chain's position when aligned on the other. Low off-diagonal PAE means the model has a definite opinion about the relative placement of the two chains.
ipTM summarizes the predicted accuracy of the interface. It is a useful ranking signal and a poor absolute one.
Neither is an affinity. A high ipTM says the model is confident in a geometry, not that the complex is tight, specific or real. Ranking designs by ipTM alone is one of the most common errors in binder pipelines.
Where it works and where it does not
Co-folding does well on complexes that resemble what the model saw in training: interfaces built from regular secondary structure, complexes between natural partners, and cases with deep sequence alignments for both chains.
It does much less well on antibodies and their antigens. The paratope sits on hypervariable loops with no useful evolutionary signal, CDR-H3 conformations are diverse, and the interface is dominated by loop packing. Published benchmarks have consistently found antibody-antigen complexes to be the weakest category, and in practice a predicted epitope should be treated as a hypothesis until an experiment supports it.
Other known weak points are complexes that require a conformational change on binding, interfaces mediated by ions, cofactors or lipids that the model was not given, and anything where the biologically relevant state is not the lowest-energy one.
Using it well
- Rank within a design campaign, not across projects. Score distributions shift with target, chain lengths and MSA depth, so a threshold that worked on one antigen is not portable.
- Look at the models, not only the numbers. Five seeds that agree on a pose is a different result from five that disagree and happen to share a score.
- Check the interface, not the fold. A design can fold beautifully and dock somewhere irrelevant.
- Follow with simulation for the shortlist. A short molecular dynamics run will tell you whether the pose survives contact with water and temperature.
- Confirm the epitope experimentally when the program depends on it: competition with a known binder, mutational scanning or a structural method.
Where that leaves it
Co-folding has made it cheap to generate plausible complexes, and cheap plausibility is not the same as evidence. It earns its place as a triage step: it removes designs that cannot be posed at all, it flags where a binder is likely to sit, and it tells you which candidates deserve the bench. Decisions belong downstream, with data.