STANFORD, 8 AUG 2026 — Sixteen AI-designed viruses infect and kill E. coli. They came out of 285 that were physically built, and the 16 that worked are not near-copies of anything in nature. Most of the coverage this week carried the first number and neither of the others.
The work was published in Science on 6 August by Brian Hie's group at Stanford University and the Arc Institute, with Stanford bioengineering graduate student Samuel King co-leading it. It first appeared as a preprint in September 2025, which is why some of the most detailed accounts of the method are nearly a year older than this week's news.
Two readings of it are circulating and both are wrong in opposite directions. One says a model can now write viruses. The other, the reflexive debunk, says the successes were just copies of the template. The published figures do not support either.
The funnel, as far as it is documented
Two genome language models, fine-tuned versions of Evo 1 and Evo 2, were prompted to generate whole genomes modelled on ΦX174 — a bacteriophage of 5,386 nucleotides encoding 11 genes, a laboratory workhorse for decades. Researchers selected 302 designs for synthesis, successfully built 285, and 16 of those turned out to be viable phages.
The public record gets vague on how many candidates the models generated before that selection. The number matters: it is the denominator that decides whether 16 successes is impressive. This week's coverage puts the figure at roughly 700,000; the Arc Institute's own account of the preprint says only "thousands". We cannot reconcile those, so we will not average them. The funnel is wide, but by how much is currently impossible to establish.
A human chose the 302
The summaries drop a filter in that funnel entirely: a human one. Whatever the number of candidates, researchers chose the 302 that went to synthesis. Nothing in the published accounts suggests the models ranked their own output and picked the winners.
That step is doing unmeasured work. If expert judgement is what separates a promising design from the rest, the capability on display is a partnership rather than an autonomous pipeline, and the honest description of the system includes the people. It also means the 5.6% is a hit rate measured after human curation, which makes it optimistic rather than conservative. The rate for an unfiltered sample is not reported.
The survivors are novel, and that is the researchers' own claim
The easy debunk fails on the facts. The functional genomes carried between 67 and 392 novel mutations relative to their nearest natural relative. One of them, Evo-Φ2147, sits at 93.0% average nucleotide identity to its closest known phage — 392 mutations away from anything sequenced. The team describes the output as "16 complete, functional, and evolutionarily novel bacteriophage genomes," and that adjective is load-bearing.
Structure backs it up. Cryo-electron microscopy of one generated phage, Evo-Φ36, found it using a shorter version of a capsid protein — 25 amino acids against 38 — sitting in a distinct orientation. That is not a typo in a copy. It is a different structural solution, reached by a model that had never seen this one.
And they work. In growth competitions and lysis kinetics, multiple generated phages showed increased fitness against ΦX174 itself. Some of the new designs beat the original.
So what does the identity gradient actually tell us
One figure in this week's reporting says viability ran at 46% among designs with at least 98% sequence identity to the template, against 5.6% across everything built. We could not corroborate that figure anywhere else, and it is the single most quoted piece of evidence for the "just copies" reading, so treat it as one outlet's number until someone else publishes it.
Even taken at face value, it says something narrower than the debunk wants. A viability rate that rises with similarity is what you would expect from any generative process working against a functional constraint: the closer you stay to a working design, the more often you get one. This describes the failure distribution, not the successes. And the successes are where the 67-to-392-mutation range and the altered capsid protein live.
Both things are true at once, which is the part that resists a headline. Most of what the models produced did not work, and what did work was genuinely new.
What the researchers ruled out on purpose
The biosafety framing is design rather than caveat. Evo cannot generate human viral sequences, because viruses that infect eukaryotes were deliberately excluded from its training data. The template was a non-pathogenic phage, the hosts were non-pathogenic laboratory strains of E. coli, and the experiments ran in dedicated biosafety cabinets with specialised disposal.
Those choices explain why this is a bacteriophage study, not a claim about human pathogens. A model trained without eukaryotic viruses has demonstrated nothing about designing one. What it has demonstrated is that the general method works at genome scale, and that is the part that will not stay confined to phages indefinitely.
The companion editorial is blunter than the paper
In a companion editorial, Thomas Inglesby and Moritz Hanke of the Johns Hopkins Center for Health Security put their conclusion in one sentence: "the ability to compose viral genomes using generative AI now exists; the governance to safely steer it does not."
The gap is specific. The practical control point in synthetic biology has been the synthesis order — screening the sequences commercial DNA providers are asked to make against databases of known hazards. Screening by matching against known agents does exactly what it was built to do, and it was not built for sequences carrying 392 mutations that no organism has ever carried. The novelty finding above is precisely what makes the screening finding urgent; they are the same fact seen from two directions.
Hsu Li Yang of the Asia Centre for Health Security in Singapore and Tom Ellis of Imperial College London are among those quoted in the coverage on where that leaves oversight. The region has a direct, if unglamorous, interest. Synthesis capacity and biotech manufacturing are distributed across Asia, and a screening regime is only as strong as its least-covered provider.
The upside is not hypothetical either
Bacteriophages kill bacteria, including bacteria that no longer respond to antibiotics, and phage therapy has been held back partly by how slow it is to find or engineer a phage that works against a particular strain.
The study's most practical result speaks to that directly. Cocktails of generated phages overcame resistance in all three ΦX174-resistant E. coli strains they were tested against, within one to five passages, where ΦX174 alone failed completely. Generating a diverse set of working variants quickly is exactly what phage therapy has needed, and this method demonstrably does it.
The same week, in the other direction
One day after the paper, Anthropic loosened the biology restrictions on its most capable model, cutting biology-related fallbacks by about 85% while keeping virology, toxicology and molecular design behind a downgrade. We covered that separately: Anthropic Will Now Answer More of Your Biology Questions.
The coincidence is instructive rather than ironic. General-purpose assistants are being opened up for health and educational questions in the same week that purpose-built genome models demonstrate design capability well beyond what those assistants will discuss. The restriction that matters is not on the chat interface.
What to watch
Whether anyone publishes the candidate count, because the width of the funnel is the difference between a promising method and a lottery. Whether the 46% identity-versus-viability figure is corroborated or quietly drops out of the record. Whether synthesis screening moves from matching known hazards to assessing novel sequences, which is the control the Johns Hopkins editorial says is missing. And whether any regulator treats genome language models as a distinct category rather than folding them into general AI oversight, where the mechanism on offer this week in Washington was a 30-day look at a chat model.