Note: This is a research note supplementing the book Unscarcity, now available for purchase. These notes expand on concepts from the main text. Start here or get the book.
Blocklists Fail When Machines Write the Unknown
Every screening system in the world is a list of things that already exist. That was fine, right up until the cost of making new things went to nearly zero.
Sixteen Things That Had Never Existed
On August 6, 2026, Science published a paper from Brian Hie’s lab at Stanford and the Arc Institute with a title so flat it undersells itself: “Generative design of bacteriophages with genome language models.” Two genome language models, Evo 1 and Evo 2, were pointed at the genome of ΦX174, a well-studied virus that eats E. coli and packs its entire existence into 5,386 nucleotides and eleven genes. The models were fine-tuned on 14,466 Microviridae sequences and asked to write more.
They wrote a great many. The team filtered down to 302 candidate genomes, successfully synthesized 285 of them, and 16 came to life: infectious bacteriophages that hunted and killed E. coli. Several worked better than the natural template. One, Evo-Φ69, outcompeted wild-type ΦX174 until it reached 65 times its starting concentration. A cocktail of the generated phages cleared bacterial strains that had already evolved resistance to the original, which is a genuinely exciting result if you care about antibiotic resistance, and you should.
Now the sentence that should have led every write-up. The viable designs carried between 67 and 392 novel mutations relative to anything in the training data, landing at 93.0% to 98.8% nucleotide identity with known sequences. Several fell below the 95% line that taxonomists use to declare a new species. These sixteen organisms did not exist. Not in a freezer, not in a database, not in the ocean. They were written.
Hold that next to how we guard this technology, and the whole architecture makes a small, embarrassed noise.
The Gate Is a Library Card
You cannot order a genome from a catalog and have it appear. Somebody has to synthesize the DNA. That step is expensive, industrial, and concentrated among a modest number of providers, which makes it the single best chokepoint in all of biosecurity. The entire defensive strategy rests on it.
And what happens at the chokepoint is this: the provider takes your requested sequence and compares it against a reference set of sequences of concern, meaning the genomes of known pathogens and the genes for known toxins. Close match, order flagged. No match, order shipped.
That is a library card check. It works exactly as well as the library is complete.
Two things make this worse than it sounds. First, in the United States the check is not legally required. The 2024 Framework for Nucleic Acid Synthesis Screening leaned on federal funding as leverage rather than on statute, and it was rescinded by Executive Order 14292 in May 2025, which instructed the Office of Science and Technology Policy to revise or replace it. As of mid-2026 no finalized replacement has been confirmed. The gate is currently held open by professional norms and an industry consortium’s good manners. Globally it is thinner still: a 2026 review drawing on the IBBIS Global DNA Synthesis Map found that only a small fraction of the 700-plus known providers worldwide maintain publicly identifiable screening at all.
Second, and more fundamentally: even a perfectly enforced, universally adopted version of this check is aimed at the wrong target. In a companion Perspective in the same issue of Science, Thomas Inglesby and Moritz Hanke of the Johns Hopkins Center for Health Security put the mismatch in one line: “The ability to compose viral genomes using generative AI now exists; the governance to safely steer it does not.” They argue that screening should be a legal obligation rather than a voluntary practice, which is plainly correct and also, on its own, insufficient. Their second point is the load-bearing one: screening methods capable of flagging genuinely new designs need to be developed, urgently, because the current ones cannot.
A machine-written genome is, by construction, the one category of thing a catalog of nature’s output does not contain.
The Catalog Assumption
Strip the biology away and you find a bet that almost every safety system in modern life is quietly making.
Call it the catalog assumption: that the set of dangerous artifacts is finite, mostly already known, and cheap to enumerate compared to the cost of creating a new one. Under that assumption, screening artifacts against a list is not just reasonable, it is the efficient answer. You spend a little on maintaining the list and you catch nearly everything, because nearly everything that shows up is a copy of something you have already seen. Novelty is rare because novelty is expensive.
Generative models are a machine for making novelty cheap. That is the entire product, and the price of running one has been in freefall for four years. The moment the cost of producing a new artifact falls below the cost of adding one to the catalog, the arithmetic inverts and the list becomes a formality: a thing you maintain, publish, cite in your compliance documentation, and quietly stop relying on.
This is the same collapse the book keeps tracking under other names. In The Verifiability Premium the argument was that professions organized around expensive verification lose their footing when a checker gets cheap. In Value Migration it was that when an input goes free, the money relocates to whatever complement stayed scarce. Screening regimes are the governance version of the same physics. When generation collapses, control migrates upstream, whether or not anyone has planned for it.
Somebody Already Ran the Experiment
You do not have to take this as theory. Eleven months before the phage paper, a team led by Microsoft’s chief scientist Eric Horvitz published the direct test in Science. They took 72 proteins of concern, including ricin and botulinum neurotoxin, and used freely available AI protein-design tools to “paraphrase” them: 76,089 variants with substantially rewritten amino acid sequences but preserved structure and active sites. Same function, different spelling.
Biosecurity screening software caught nearly all the originals. Many of the paraphrases walked straight through. The team did the responsible thing, developing patches with the screening vendors and the International Gene Synthesis Consortium before publishing. After patching, the tools still missed roughly 3% of variants.
Three percent sounds like a rounding error until you remember what the denominator is. This is not a system where 97% coverage means 97% safety; it means an adversary generating candidates in bulk gets a large number of tickets. And the patch was tuned against the paraphrasers that existed in 2025. Michael Cohen of UC Berkeley named the uncomfortable part out loud in MIT Technology Review: “There seems to be an unwillingness to admit that sometime soon, we’re going to have to retreat from this supposed choke point.” Adam Clore of Integrated DNA Technologies, one of the companies doing the screening, was blunter about the maintenance burden: “The patch is incomplete, and the state of the art is changing. But this isn’t a one-and-done thing.”
That is what defending a catalog against a generator looks like from the inside. You are not solving a problem. You are subscribing to one.
This Is Not a Biology Story
If reference-matching only failed in biosecurity, it would be a narrow technical problem for a narrow technical community. It fails everywhere it is deployed, and the failures rhyme so precisely that you can predict them.
Malware. Antivirus began as signature matching, which is reference-matching with a different vocabulary, and it lost this fight thirty years ago to polymorphic code. The industry rebuilt around behavioral detection, sandboxing, and reputation. Then the generation cost fell again. On November 5, 2025, Google’s Threat Intelligence Group described PROMPTFLUX, a VBScript dropper that calls the Gemini API roughly once an hour with a request to rewrite its own source code to evade antivirus signature detection, then saves the fresh variant to disk. It was caught in a development state and never demonstrated the ability to compromise a network. That is not the point. The point is that a piece of malware now ships with a subscription to a novelty generator, and GTIG’s recommendation to defenders was to stop leaning on signatures.
Identity and fraud. Know-your-customer systems, document verification, and fraud models are trained on catalogs of known-bad patterns, known-bad documents, known-bad faces. Deepfake fraud now accounts for 6.5% of all fraud attempts globally, up from 0.1% in 2022, per Sumsub’s identity-fraud data. Gartner called the endgame two years early: by 2026, deepfake attacks on face biometrics would mean 30% of enterprises no longer consider identity verification and authentication solutions reliable in isolation. That is an entire industry publicly downgrading its own catalog.
Three domains, three different sets of experts, one architecture, one failure mode. Biosecurity is not experiencing a novel crisis. It is arriving, late and with much higher stakes, at the place antivirus reached in 1995.
“But the Hit Rate Was 16 out of 302”
The strongest objection deserves a straight answer. Jordi García Ojalvo of Pompeu Fabra University made it to the Science Media Centre: out of thousands of generated genomes only 16 were viable, each design had to be built and tested in a lab one at a time, and it is hard to imagine these models producing working genomes out of the box. A 5% success rate demanding a wet lab, months of work, and real expertise is not a garage-scale threat. Correct on every count.
It is also the wrong variable to watch. Hit rates are the single most reliable thing to improve in machine learning; that is what the last decade has consisted of. And the argument here does not actually depend on the hit rate at all. It depends on one asymmetry: the reference set is finite and grows by human curation, while the generated set is unbounded and grows by compute. Even at a 1% success rate, a generator that can propose a million candidates has decoupled itself from any list a human maintains. The defense is not “novelty is hard.” The defense was always “novelty is expensive,” and that is the specific sentence that stopped being true.
Note too that the phage team’s own safeguard was a blocklist one level up. Evo’s training data deliberately excluded viruses that infect humans, animals, and plants, so the model cannot write what it never read. That is a real and commendable control, and it is also an enumeration: it works against the categories someone thought to remove. Every layer of this defense is the same shape.
What Actually Replaces It
The move is not a longer list. It is a different question. Instead of asking have I seen this artifact before, you ask things that a generator cannot make cheap:
Gate the function, not the sequence. In May 2026 a consortium of 30 authors published “Beyond sequence similarity: toward function-based screening of nucleic acid synthesis”, arguing for screens that detect “sequences whose predicted molecular properties indicate a capacity for biological functions of concern” regardless of what they resemble. Biophysics constrains what a functional toxin can be in a way that spelling does not. Use the same models to predict capability that the adversary is using to generate it.
Gate the capability. Synthesizers, model weights, and the compute to run them are physical, expensive, and countable. They are the layer that stays scarce when the artifacts stop being scarce, which is precisely why control is migrating there whether we design for it or not.
Gate provenance. Ask not what a thing is but where it came from and who will vouch for it. This is the verifiability inversion applied to security, and the cryptographic machinery for doing it without building a surveillance apparatus already exists in zero-knowledge proofs: prove the attribute, reveal nothing else.
Gate the customer, and attach liability. Know who is ordering and make somebody legally answerable when the order was wrong. As The Liability Gap argues about professional licensing, an accountable human is a different kind of safeguard from a technical one, and it degrades far more gracefully. Screening the customer is not glamorous. It also does not care how novel the sequence is.
Every one of these moves control upstream, away from the artifact and toward the pipe. Which brings the honest problem into view.
The Unscarcity Reading
Upstream is where the chokepoints live, and chokepoints have owners.
Shift security from screening outputs to gating capability, and you have handed enormous discretionary power to a handful of synthesis providers, model labs, and cloud operators. Nobody elected them. This is the exact fork Value Migration keeps landing on, arriving now in a domain where the downside is not a bad quarter: when a scarce layer becomes the control point for everyone downstream, is it run as a utility with published rules and an appeals process, or as a private toll road priced to whoever needs it most? The physics of generative biology does not answer that. Governance does, and right now the American answer is a rescinded framework and a replacement that has not shipped, while the capability it was written to govern went to press in Science.
That gap is the book’s argument in miniature. Institutions do not fail because the people inside them are foolish. They fail because they were engineered against a cost curve that no longer holds, and cost curves move faster than statutes. The 2024 framework was a competent piece of work aimed at a world where writing a new genome was harder than looking one up. That world ended on a Thursday in August.
The uncomfortable part is that none of this argues for stopping. Sixteen designed phages that beat antibiotic-resistant bacteria is the good outcome, the Foundation kind of outcome, the reason abundance in biology is worth building toward at all. Suppressing the capability would cost more lives than it saved and, as the governance literature keeps finding, there is no mechanism that could enforce it anyway. What it argues for is dropping the pretense that a list of yesterday’s dangers is a security architecture, and paying for the harder thing: screens that reason about function, chokepoints run as accountable utilities, provenance you can verify, and a named party who answers when it goes wrong.
When the printing press collapsed the marginal cost of a page, the institutional response was a list: the Index Librorum Prohibitorum, an enumeration of forbidden titles, first issued in 1559. It was the best available tool and it was already obsolete on arrival, because presses could produce new titles faster than committees could name them, and appearing on the list turned out to be excellent publicity. Five centuries later we are running the same play against a machine that writes genomes. Shorter fuse. Considerably worse failure mode.
Related Articles
- The Verifiability Premium - When checking gets cheap, the whole trust apparatus reorganizes around proof
- Value Migration - Where control goes when the input collapses to zero
- The Commoditization of Intelligence - The generation-cost collapse driving all of this
- Zero-Knowledge Proofs - Verifying provenance without building a surveillance state
- The Liability Gap - Why an accountable human is a different safeguard than a technical one
- When the Lab Pulls Its Own Alarm - What happens when capability outruns the machinery meant to steer it
- Goodhart’s Law and AI Governance - What a proxy measure does once someone optimizes against it
- Failed Architectures of Law - Why exhaustive rulebooks lose to adaptable judgment
Every institution you rely on is a standing bet about what stays expensive. This is a book about what happens when those bets stop paying, and what to build before they do. Read the book or start with the preamble.