Protein language models
A protein language model is trained on large numbers of protein sequences, usually by masking residues and learning to predict them. Nothing about structure or function is supplied. What the model ends up with is a statistical picture of what protein sequences look like, expressed as per-residue and per-sequence vectors, plus the ability to score how likely a given residue is in a given context.
That turns out to be useful for a narrow, real set of jobs, and it is routinely oversold for a wider set.
What they are good at
Representations for downstream models. Embeddings are a strong starting point for supervised models with limited labeled data, because they encode a lot of what is conserved in protein sequences. This is where most of the practical value sits.
Fitness and variant effect prediction. Likelihood under the model correlates with tolerance to mutation in many proteins. Scanning variants for which ones the model considers unusual is a reasonable filter.
Structure prediction without alignments. Single-sequence structure models built on language models trade accuracy for speed and for independence from sequence search, which matters when you need to fold very many sequences.
Generation of plausible sequences. Sampling produces sequences that look like proteins. Whether they fold and do what you want is a separate question that needs a separate check.
Where they are weak
Antibodies. The evolutionary signal that makes these models work is mostly absent in CDRs, which are diversified somatically rather than over evolutionary time. General models tend to be least informative exactly where an antibody program needs help. Antibody-specific models trained on repertoire data help, and they inherit the biases of the repertoires they were trained on.
Binding. A language model has no notion of a partner. Likelihood is not affinity, and a high-likelihood sequence is not a good binder.
Designed proteins. A de novo sequence that is nothing like anything in the training data will score poorly whether or not it works. Low likelihood is not evidence of failure for designed molecules.
Reading claims about them
Two questions separate a useful benchmark from a misleading one. First, what was held out? Sequence identity splits are weak, because homologs share more than sequence identity suggests; cluster-level and temporal splits are stronger. Second, what is the baseline? A simple conservation score or a model trained on a few obvious features often performs close to a large model on the same task, and knowing that changes how much weight the result deserves.
Using them in a campaign
Put plainly, these models are good filters and poor deciders. They remove sequences that look wrong, they encode useful priors for a supervised model when data are scarce, and they generate diversity worth testing. They do not replace a measurement, and an unvalidated model-derived score does not belong in a decision that costs real money.