Biomolecular Modeling · AI for Science · Language Models

AlphaFold 3 vs ESM3 vs ProGen2: Which Protein Design Model for Which Job?

Protein design is five different tasks wearing one name: AlphaFold 3 predicts structures, ESM3 and ProGen2 generate sequences, DynamicMPNN does inverse folding, and Feynman-Kac steering drives a diffusion designer.

AlphaFold 3 vs ESM3 vs ProGen2: Which Protein Design Model for Which Job?

The one split: prediction, generation, and control are different jobs

“Protein design” sounds like one problem. It is at least three. Predicting how a given sequence folds is a regression problem with a known answer. Generating a sequence that folds into a desired shape, or that carries a desired function, is a search problem over an astronomical space. Steering an already powerful generator toward a measurable property is a control problem on top of that. The five papers in this comparison each own one of these jobs, and almost every confused take about “which protein AI is best” comes from scoring them against each other as if they were one leaderboard.

The practical consequence: the numbers in the table below sit in different units, on different benchmarks, against different baselines. A guide that ranked them 1 to 5 would be inventing a comparison. What you can do, and what this page does, is map each model to the task it was built for, quote the number each paper actually reports, and mark the rows where even that is only a same-task, same-baseline claim.

The five models in one paragraph each

AlphaFold 3 is a structure predictor, not a designer in the generative sense. You hand it a mixed system (protein chains, DNA, RNA, ligands, ions, modifications) and it returns the joint 3D structure, using a diffusion network that generates raw atom coordinates directly. Its contribution is breadth: one model covers interaction types that used to need a half-dozen specialized tools.

ESM3 is a generator. It models protein sequence, structure, and function as three coordinated token tracks, so you can prompt it with a structural motif or a functional constraint and let it fill in a sequence. Its landmark result is a generated fluorescent protein only 58% identical to the nearest known fluorescent protein, which the authors estimate is an evolutionary distance of over 500 million years, and which was synthesized and found to fluoresce.

ProGen2 is also a generator, but its question is about scale and data rather than prompting. The suite goes up to 6.4B parameters, trained on sequence datasets drawn from over a billion proteins spanning genomic, metagenomic, and immune repertoire databases. The paper claims state-of-the-art results on capturing the distribution of observed evolutionary sequences, generating novel viable sequences, and predicting protein fitness with no finetuning. Notably, it releases models and code.

DynamicMPNN is an inverse folding model: backbone geometry in, sequence out. Its twist is that real proteins switch between conformations, so it trains jointly across conformational ensembles instead of averaging single-state predictions afterward. It was trained on 46,033 conformational pairs covering 75% of CATH superfamilies, and evaluated by folding its designed sequences with AlphaFold 3.

Feynman-Kac steering is not a designer at all; it is a steering layer for RFdiffusion, the established diffusion protein designer. It builds guiding potentials from ProteinMPNN sequence recovery plus structural relaxation, resamples trajectories during denoising according to a user-defined reward, and reports increasing binder designability by 89.5% while working with objectives that are not differentiable.

Key numbers

ModelInput to outputScale and training dataHeadline numberBaseline and settingSame task as the others?
AlphaFold 3Mixed biomolecular system to predicted 3D complexDiffusion over raw atom coordinates; no ligand-specific scaffoldBeats classical docking without the binding pocket given; beats AF-Multimer v2.3 on antibody-antigenDocking tools, nucleic-acid predictors, prior DeepMind systemsNo: prediction, not generation
ESM3Prompts over sequence, structure, function tracks to sequence1.4B, 7B, and 98B parameter modelsFluorescent protein at 58% identity to nearest known; estimated 500M plus years of evolutionary distanceSingle generated protein, wet-lab validated for fluorescencePartially: generation, but n of 1 function
ProGen2Sequence to sequenceUp to 6.4B parameters; over 1B proteins from genomic, metagenomic, and immune repertoire dataState-of-the-art claimed on 3 axes: evolutionary distribution, novel viable sequences, zero-shot fitnessPrior protein language models, per-paper benchmarksNo: claims cover three different evaluations
DynamicMPNNConformational ensemble to sequenceTrained on 46,033 conformational pairs covering 75% of CATH superfamiliesUp to 25% better decoy-normalized RMSD and 12% better sequence recovery than ProteinMPNNProteinMPNN, on a multi-state benchmark, evaluated with AlphaFold 3Yes for inverse folding, but multi-state only
Feynman-Kac steeringRFdiffusion trajectory plus reward to steered designGuiding potentials from ProteinMPNN and structural relaxationBinder designability up 89.5%Unsteered RFdiffusion, predicted interface energetics also improveNo: a control method, not a designer

Read the table for what it refuses to do. There is no column where two models share a benchmark and a metric, so no ranking can be computed from these rows. The three quantitative rows that do compare against a baseline (DynamicMPNN against ProteinMPNN, Feynman-Kac against unsteered RFdiffusion, and ESM3’s single validated protein) each say something real but narrow: an improvement inside one paper’s own protocol. Treat each number as “this paper, this setting,” never as a property of the model in general.

Two honesty notes belong here. First, AlphaFold 3’s headline is qualitative by design: the paper’s value is one system beating specialized tools across many interaction types, and it deliberately offers no single trophy score. Second, ESM3’s 58% figure is the distance of one designed protein to its nearest natural relative, not a success rate. Both are easy to misread as leaderboard numbers, and neither is.

When to use which

  • You have sequences and need structures or complexes: AlphaFold 3, through the AlphaFold Server if you do not want to run it yourself. Remember it predicts a static most-likely conformation, not binding free energy or kinetics, and its full weights were not opened the way AlphaFold 2’s were.
  • You want to generate sequences under a structural or functional constraint: ESM3. The three-track prompting is the point. Read the 58% fluorescent protein as a demonstration of reach, not as an average success rate, and remember that fluorescence is one function.
  • You need to rank or score variants, or generate at scale without task finetuning: ProGen2. The 6.4B model trained on over a billion sequences is aimed exactly at zero-shot fitness and broad sequence modeling, and it ships code. Validate on your own assay before trusting the zero-shot claim.
  • Your target protein has multiple functional conformations: DynamicMPNN. Single-state inverse folding optimizes one backbone; if the biology lives in the switching between states, the multi-state training data (46,033 pairs, 75% of CATH superfamilies) is the relevant difference.
  • You already use RFdiffusion and want a property it was not trained for: Feynman-Kac steering. The reward can be anything you can score, differentiable or not, which is exactly the gap it was built to close.

Limits and open questions

The deepest limit is shared: almost none of these numbers survive contact with a different lab’s assay. A designability rate, a sequence recovery percentage, and a docking accuracy are answers to different questions, and each is measured under its own protocol. When you move from reading these papers to spending wet-lab budget, the paper numbers function as priors for which method to try first, not as predictions of your success rate.

Reproducibility is uneven across the five. ProGen2 releases models and code. AlphaFold 3 at release kept its weights behind the AlphaFold Server with usage limits, which constrained independent scrutiny. ESM3’s full 98B scale is not something most labs can verify or run. DynamicMPNN and Feynman-Kac sit on top of other people’s systems (ProteinMPNN, RFdiffusion, AlphaFold 3), so their results inherit those systems’ failure modes.

The open scientific question is evaluation. Protein design has no shared, cheap, high-throughput ground truth. Fluorescence worked for ESM3 because it is one function with a one-readout assay. Binder designability, as in the Feynman-Kac paper, is judged by predicted energetics and foldability, which are themselves model outputs. Until experimental throughput catches up with generation speed, every headline number in this field should be read with that chain of surrogates in mind.

FAQ

Is AlphaFold 3 a protein design model?

Strictly, no: it is a structure prediction model, and its job is the inverse of design. You give it a sequence (plus ligands, nucleic acids, and so on) and it gives you the 3D structure. Designers like ESM3 or RFdiffusion go the other direction, from a desired structure or property to a sequence. In practice the two are chained constantly, as DynamicMPNN’s use of AlphaFold 3 as its evaluator shows.

What is the difference between ESM3 and ProGen2?

Both are protein language models that generate sequences, but they answer different questions. ESM3 is built for controllable generation: you prompt it across sequence, structure, and function tracks, and its signature result is a 58% identity fluorescent protein, validated in the lab. ProGen2 is built to probe scale and data: models up to 6.4B parameters trained on over a billion sequences, with state-of-the-art claims on evolutionary modeling and zero-shot fitness. If you need steering, think ESM3; if you need a fitness scorer or a broad sequence prior, think ProGen2.

How does DynamicMPNN compare to ProteinMPNN?

On its own multi-state benchmark, trained on 46,033 conformational pairs covering 75% of CATH superfamilies, DynamicMPNN reports up to 25% better decoy-normalized RMSD and 12% better sequence recovery than ProteinMPNN, with designs evaluated by AlphaFold 3. The catch is scope: this is a claim about proteins that adopt multiple conformations. For a rigid single-state target, the extra machinery buys nothing.

What does the 89.5% in the Feynman-Kac steering paper mean?

It is a relative increase in binder designability when RFdiffusion is steered with the paper’s Feynman-Kac guiding potentials, built from ProteinMPNN sequence recovery and structural relaxation. It is not an absolute designability rate, and it is not a comparison against ESM3 or any other generator; the baseline is unsteered RFdiffusion. Its real significance is that the reward can be arbitrary and non-differentiable.

Which of these models is open source?

ProGen2 explicitly releases its models and code. AlphaFold 3 does not: access runs through the free AlphaFold Server with usage limits, which drew criticism for slowing independent reproduction. ESM3 was released through EvolutionaryScale with API access rather than full open weights at the reported 98B scale. DynamicMPNN and Feynman-Kac steering build on ProteinMPNN and RFdiffusion, which have their own release terms, so check those stacks before planning a pipeline around them.