<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>point.free - biology</title>
    <subtitle>Notes from Christina Sørensen — software engineer, NixOS Steering Committee, author of eza.</subtitle>
    <link rel="self" type="application/atom+xml" href="https://point.free/tags/biology/atom.xml"/>
    <link rel="alternate" type="text/html" href="https://point.free"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-07-25T00:00:00+00:00</updated>
    <id>https://point.free/tags/biology/atom.xml</id>
    <entry xml:lang="en">
        <title>The AI bioweapon risk isn&#x27;t jailbreaks</title>
        <published>2026-07-25T00:00:00+00:00</published>
        <updated>2026-07-25T00:00:00+00:00</updated>
        
        <author>
          <name>
            
              Christina Emilie Sørensen
            
          </name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://point.free/blog/the-ai-bioweapon-risk-isnt-jailbreaks/"/>
        <id>https://point.free/blog/the-ai-bioweapon-risk-isnt-jailbreaks/</id>
        
        <content type="html" xml:base="https://point.free/blog/the-ai-bioweapon-risk-isnt-jailbreaks/">&lt;h3 id=&quot;what-a-danish-threat-assessment-gets-right-about-ai&quot;&gt;What a Danish threat assessment gets right about AI&lt;&#x2F;h3&gt;
&lt;p&gt;I read &lt;em&gt;Det biologiske trusselsbillede 2026&lt;&#x2F;em&gt;, published this summer by the &lt;a href=&quot;https:&#x2F;&#x2F;www.biosikring.dk&quot;&gt;Centre for Biosecurity and Biopreparedness (CBB)&lt;&#x2F;a&gt; at Statens Serum Institut. It’s the first update since 2020 to Denmark’s national assessment of man-made biological threats, and on artificial intelligence it is more careful than most of what gets written on the subject.&lt;&#x2F;p&gt;
&lt;p&gt;The reason to care right now is that the guardrails are already here, and
already contested. Anthropic’s recent limits on biological work have frustrated
a lot of people doing entirely legitimate research, down to the hobbyist end of
synthetic biology.&lt;&#x2F;p&gt;
&lt;p&gt;Tools like &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Binomica-Labs&#x2F;SpliceCraft&quot;&gt;SpliceCraft&lt;&#x2F;a&gt;, a local
plasmid-design and cloning workbench of the sort working labs use every day can
no longer use frontier models; and the argument over where to draw those lines
is happening mostly without good evidence about where the risk really sits.&lt;&#x2F;p&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img src=&quot;https:&#x2F;&#x2F;point.free&#x2F;blog&#x2F;the-ai-bioweapon-risk-isnt-jailbreaks&#x2F;splicecraft.png&quot; alt=&quot;SpliceCraft Screenshot&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;figcaption&gt;SpliceCraft offers an example of the real usecase for large language models in making tools that help advance lifescience. It is also a quite stunningly beautiful tool.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;Which is where a report like this earns some attention.&lt;&#x2F;p&gt;
&lt;p&gt;Denmark runs a serious bio and life-science sector, Novo Nordisk (that readers
will know for inventing Ozempic and Wegovy) and the cluster around it, and its
Statens Serum Institut is a working public-health body with real operational
responsibilities. It has largely kept its footing at a point when some
comparable institutions elsewhere have grown more politicised, or lost the
thread.&lt;&#x2F;p&gt;
&lt;p&gt;Anyways, onto the report. For actors without a professional background, the advantage is limited. For competent ones it’s significant, above all in developing new agents. A model can be built into specialised research tools, predict what a given genetic modification will do, design genes from scratch, and check a sequence for its likely effect before anyone synthesises anything.&lt;&#x2F;p&gt;
&lt;p&gt;In practice the two halves often get conflated as the same. Some guardrails seem to stop only the novice half — when they don’t just block everything — because the novice is the case you can measure, and you build the guardrail against. The thing you can score. Hand a model to someone with no background and check whether they get any further than a search engine would take them.&lt;&#x2F;p&gt;
&lt;p&gt;And when that approach fails, just block everything is the new solution. I think that’s an avoidable limitation.&lt;&#x2F;p&gt;
&lt;p&gt;CBB cites two studies. The first came out of &lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2306.03809&quot;&gt;Kevin Esvelt’s lab at MIT&lt;&#x2F;a&gt; in 2023. Students were given an hour with a chatbot, and in that hour it suggested four candidate pandemic pathogens, explained how to make them from synthetic DNA, and pointed them at synthesis firms unlikely to screen the order. The second, a 2024 &lt;a href=&quot;https:&#x2F;&#x2F;www.rand.org&#x2F;pubs&#x2F;research_reports&#x2F;RRA2977-2.html&quot;&gt;RAND red-team study&lt;&#x2F;a&gt;, went further. Teams role-playing hostile actors drew up plans for a biological attack, some with a large language model and some without, and the plans came out no more viable either way.&lt;&#x2F;p&gt;
&lt;p&gt;CBB’s own verdict on today’s chatbots is deflationary. Because they hallucinate, the output has to be checked by someone who already knows the answer, and the report expects it to be a while before this technology is much use to anyone without prior training and hands-on experience with dangerous biological material.&lt;&#x2F;p&gt;
&lt;p&gt;However, the competent-actor half is what CBB flags as consequential. Butasome vocabulary first.&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Term&lt;&#x2F;th&gt;&lt;th&gt;What it means&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Biosafety&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Protection against accidents.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Biosecurity&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Protection against deliberate misuse.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Weaponisation&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;The step between having an agent and having a weapon. Historically it has defeated almost everyone who tried it.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Dual use&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;A technique that serves a legitimate purpose and a harmful one without changing form. This is why intent is the thing regulation keeps reaching for, and keeps failing to hold.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;So where do language models actually change things? Four places, I think. I’ve ordered them the way public discussion usually ranks them, which is roughly the opposite of how much I think they matter.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-untrained-attacker&quot;&gt;The untrained attacker&lt;&#x2F;h2&gt;
&lt;p&gt;This is the scenario that gets the most attention and has the least evidence behind it.&lt;&#x2F;p&gt;
&lt;p&gt;It’s easy to see why. Tools like &lt;a href=&quot;https:&#x2F;&#x2F;vibe-genomics.replit.app&#x2F;&quot;&gt;Vibe Genomics&lt;&#x2F;a&gt;, a browser app built around the idea of doing genomics by plain-language prompting, make the leap feel short: if a model can walk you through that, why not through building something dangerous? I don’t think that intuition is stupid. I do think the reasons to be sober about it are stronger, and they have less to do with what the software can say than with how living material actually behaves.&lt;&#x2F;p&gt;
&lt;p&gt;It dominates because it’s tractable. You can build an evaluation for it, red-team it, and publish the score. The frontier labs have converged on explicit capability thresholds that trigger additional safeguards, which CBB notes approvingly, and those thresholds are largely written in terms of what an unsophisticated user can elicit.&lt;&#x2F;p&gt;
&lt;p&gt;It’s also the scenario the historical record argues against hardest. Take Aum Shinrikyo, the Japanese doomsday cult behind the 1995 sarin attack on the Tokyo subway. Before they reached for nerve gas they spent years trying to build biological weapons, with money, laboratories, and members who held real scientific degrees. The &lt;a href=&quot;https:&#x2F;&#x2F;www.cnas.org&#x2F;publications&#x2F;reports&#x2F;aum-shinrikyo-insights-into-how-terrorists-develop-biological-and-chemical-weapons&quot;&gt;most detailed account we have&lt;&#x2F;a&gt;, a reconstruction led by former US Navy Secretary Richard Danzig, counts at least six attempted biological attacks between 1990 and 1995, five with botulinum toxin and one with anthrax. Every one of them failed. They grew the wrong strains, couldn’t get the spores to dry, and never worked out how to spread what little they had. Biology beat them, so they switched to chemistry.&lt;&#x2F;p&gt;
&lt;p&gt;Islamic State is the other example. At its peak it ran chemical and biological weapons efforts with hundreds of people and a budget of more than ten million US dollars a year. It managed chemical attacks dozens of times; on the biological side, &lt;a href=&quot;https:&#x2F;&#x2F;ctc.westpoint.edu&#x2F;the-islamic-state-and-wmd-assessing-the-future-threat&#x2F;&quot;&gt;the documented record&lt;&#x2F;a&gt; shows plenty of ambition and some grim experimentation, but no weapon it could actually field.&lt;&#x2F;p&gt;
&lt;p&gt;Access to information isn’t what stopped either of them. Biological material is alive, and it behaves like it — thrown off by something as ordinary as the pH of the water it sits in, and apt to lose one property the moment you push hard on another.&lt;&#x2F;p&gt;
&lt;p&gt;The Soviet programme, the largest the world has ever seen, kept running into exactly that, by the later testimony of the scientists who ran it… &lt;strong&gt;Make a strain more lethal and it often becomes less able to spread.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;A chatbot that answers questions well does nothing about any of this. What would matter is a model that could stand in for hands-on laboratory experience, and that’s a different capability from answering questions.&lt;&#x2F;p&gt;
&lt;p&gt;I’m not saying novice uplift is impossible. My worry is narrower: if it soaks up most of the evaluation effort because it happens to be the easiest one to measure, that’s a poor reason to let it set the agenda, and it will distract from the measures that may have more potential.&lt;&#x2F;p&gt;
&lt;p&gt;None of this contradicts what comes next. The binding constraint just depends on what the actor already owns. Aum’s missing input was tacit skill and organisation — no quantity of information substitutes for knowing how to dry spores, which is why the historical record looks the way it does. For an actor who already has the lab, the staff and the hands, that constraint is already satisfied, and the next one along is design search over a space too large to explore by hand. That’s the one a model moves. So “information was never the bottleneck” and “design acceleration matters enormously” aren’t in tension; they describe different actors at different points on the same curve.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-trained-scientist&quot;&gt;The trained scientist&lt;&#x2F;h2&gt;
&lt;p&gt;This is the one CBB names, and I think it’s the one that should worry alignment people most, because the problem is structural rather than technical.&lt;&#x2F;p&gt;
&lt;p&gt;The report is specific. AI integrated into specialised research tools can substantially improve an actor’s ability to predict the outcome of genetic modification, design novel genes from first principles, and screen sequences for effect before committing to synthesis. Advanced programmes already carry misuse potential today. It is theoretically possible, CBB says, to predict novel proteins or protein variants more potent than anything currently known.&lt;&#x2F;p&gt;
&lt;p&gt;Picture the stream of questions a tool like this actually gets. A working scientist asks a hundred ordinary things: protein structure, expression systems, how to purify a sample, how to keep it stable. Every one of those questions is also asked thousands of times a day by people doing entirely legitimate work, and you can’t refuse any of them without breaking the tool for the whole field it serves. None of them is “how do I make a weapon.”&lt;&#x2F;p&gt;
&lt;p&gt;Refusal training is calibrated on intent expressed in a prompt. The competent actor doesn’t put intent in the prompt. It went into the choice of project, which happened before the conversation opened and is invisible to the model. CBB identifies the same problem in the physical regime, and is explicit about it: research, development and production of biological weapons can be hard to distinguish from peaceful, entirely legitimate commercial or research activity. What I’m describing is that same problem, moved off the laboratory bench and onto the model’s inputs.&lt;&#x2F;p&gt;
&lt;p&gt;I don’t think it’s solvable at the level of the individual query, and better classifiers won’t change that.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-screening-problem&quot;&gt;The screening problem&lt;&#x2F;h2&gt;
&lt;p&gt;Here AI works both sides of the problem at once. It sharpens the attack, and it’s meant to sharpen the defence too, but the defence is the side that’s behind.&lt;&#x2F;p&gt;
&lt;p&gt;Commercial gene synthesis is the chokepoint the whole non-proliferation story rests on, and as chokepoints go it’s a reasonable one. A handful of firms take DNA orders, and they check those orders against a list of sequences known to be dangerous. CBB is blunt about the weak spots: taking part is voluntary, and the screening tools have already proven leaky against orders that were deliberately suspicious.&lt;&#x2F;p&gt;
&lt;p&gt;In October 2025 a team led by &lt;a href=&quot;https:&#x2F;&#x2F;www.science.org&#x2F;doi&#x2F;10.1126&#x2F;science.adu8578&quot;&gt;Eric Horvitz at Microsoft showed in &lt;em&gt;Science&lt;&#x2F;em&gt;&lt;&#x2F;a&gt; how much worse it can get. They took known toxins and used open-source protein-design models to generate tens of thousands of redesigned versions, different enough in their DNA to slip past the screening software but predicted to keep working as toxins. To their credit, they held the finding back until fixes had been written and handed to the screening providers first, borrowing the coordinated-disclosure habit from computer security. It’s the rare bit of good news here. It doesn’t touch the underlying problem, though: the screening works by matching new orders against known dangerous sequences, and design tools now produce dangerous sequences that match nothing on file.&lt;&#x2F;p&gt;
&lt;p&gt;Two developments compound it.&lt;&#x2F;p&gt;
&lt;p&gt;The first is hardware. Enzymatic benchtop synthesisers, machines that print DNA to order on the lab bench, are expected to reach around seven kilobases within a few years. DNA is measured in base pairs, the individual letters of the genetic code, and a kilobase is a thousand of them; seven kilobases is a run about seven thousand letters long. The longer the fragments a machine can print, the fewer you need to assemble a full genome, and the more of the work moves off a screened commercial service and onto an unscreened instrument in the lab.&lt;&#x2F;p&gt;
&lt;p&gt;The second is horsepox. In work published in 2018, a team under &lt;a href=&quot;https:&#x2F;&#x2F;journals.plos.org&#x2F;plosone&#x2F;article?id=10.1371%2Fjournal.pone.0188453&quot;&gt;David Evans at the University of Alberta&lt;&#x2F;a&gt; rebuilt the virus, a relative of the one that causes smallpox that most people had assumed was gone for good, by stitching together ten DNA fragments ordered by mail from a commercial supplier. It reportedly cost something like 100,000 US dollars. Their stated goal was a safer smallpox vaccine, and publishing the method touched off a long argument about whether it should have appeared in print at all. The finished genome ran to 212,000 base pairs, and it settled the question of whether the mail-order route works.&lt;&#x2F;p&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img src=&quot;https:&#x2F;&#x2F;point.free&#x2F;blog&#x2F;the-ai-bioweapon-risk-isnt-jailbreaks&#x2F;horsepox-virus.webp&quot; alt=&quot;Horsepox Virus graphic&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;figcaption&gt;Graphic from the paper on Horsepox Vaccine synthesis.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;Which makes this not really a misuse scenario. Nobody has to jailbreak anything. Legitimate protein design capability, deployed exactly as intended, degrades a control system built on the assumption that dangerous sequences resemble known dangerous sequences.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-team-barrier&quot;&gt;The team barrier&lt;&#x2F;h2&gt;
&lt;p&gt;This is the one I think is underrated, and my reasons come from history more than from anything about the models themselves.&lt;&#x2F;p&gt;
&lt;p&gt;The American programme, until Nixon shut it down in 1969, employed microbiologists, ornithologists, veterinarians, engineers, chemists, airborne-dispersal specialists and statisticians. That list is the real barrier, and what it describes is organisation. A lone actor fails because no single person is a whole research team, and putting a team together is exactly what regulators and intelligence services are good at disrupting.&lt;&#x2F;p&gt;
&lt;p&gt;CBB has a term for that disruption, operative pressure: the work of making materials, equipment and expertise hard to come by, so that anyone hostile has to work covertly, under unstable and less-than-ideal conditions. Aum ran into this directly. Its programme was interrupted again and again, its agents and equipment destroyed, because the cult kept fearing the police were about to close in.&lt;&#x2F;p&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img src=&quot;https:&#x2F;&#x2F;point.free&#x2F;blog&#x2F;the-ai-bioweapon-risk-isnt-jailbreaks&#x2F;ct-pressure.webp&quot; alt=&quot;Counter Terrorist Pressure&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;figcaption&gt;The US CTC graphic on Counter-Terrorist (CT) Pressure.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;A model that could stand in for the team is where this defence stops working. The team was the thing being disrupted, and a single person has nothing to assemble.&lt;&#x2F;p&gt;
&lt;p&gt;There’s a second mechanism. Operative pressure is external; it makes acquisition hard. But there has always been an internal deterrent running alongside. The professional class capable of constituting a material threat actor has a career, a standing, a lab, a future. That stake has been doing quiet work no regulation is doing. CBB estimates up to 30,000 people worldwide hold a virology doctorate qualifying them for advanced work with viruses, and notes the pool keeps growing as more high-security laboratories are built and staffed.&lt;&#x2F;p&gt;
&lt;p&gt;It isn’t uniformly invested. Suppose automation-driven displacement in biotechnology and pharmaceutical research turns out to be real, meaning genuine capability substitution rather than shareholder-pleasing restructuring. Then for the first time the population of people who have the training and have lost the stake grows as a direct consequence of the same technology.&lt;&#x2F;p&gt;
&lt;p&gt;The nearer-term version has nothing to do with lone actors. It’s recruitment. Islamic State attempted to recruit scientists and technicians out of the United States, France, the United Kingdom and Belgium. Organisations that already have structure and money have always been shopping for exactly this expertise, and a displaced professional is far more plausibly recruited into an existing programme than spontaneously becoming one.&lt;&#x2F;p&gt;
&lt;p&gt;I should be clear that this is forecasting rather than observation. Current models don’t substitute for a research team, and CBB’s assessment is that they aren’t close. Whether the labour effect materialises at scale is unknown to me. But it’s the scenario where the alignment question and the economic question turn out to be the same question.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;CBB’s summary line puts it plainly: risk depends increasingly on intention rather than on technological and accessibility limitations.&lt;&#x2F;p&gt;
&lt;p&gt;That sentence should be uncomfortable for anyone doing safety evaluation. If we
only measure what a model will say to a user who reveals hostile intent in the
asking, we leave three of the four scenarios above on the table. If we naively block everything, we bottleneck scientific advancement on some of the most pressing scientific issues.&lt;&#x2F;p&gt;
&lt;p&gt;The competent actor asks legitimate questions; the screening problem arises from
tools working correctly; the labour effect is a second-order consequence of
deployment succeeding. Filling the gap between legitimate and illegitimate use
is much more than a blanket ban or mere intentions.&lt;&#x2F;p&gt;
&lt;p&gt;To me the gap struggles to be filled because of a general lack of
consequentialism in both the models’ training and their output. The attempt at
“unbiased” models really just produces unopinionated ones, which end up deciding
on intentionality or all-or-nothing rather than on what outcomes a given piece of
information and work will lead to.&lt;&#x2F;p&gt;
&lt;p&gt;The essay has already established that the thing that actually moves the needle
is the long directed search — design search over a space too large to explore by
hand: “that’s the one a model moves”.&lt;&#x2F;p&gt;
&lt;p&gt;There’s no magic unlock where a single answer conjures a weapon from nothing. So
if the real capability is a long, shaped trajectory of work, then the real
intervention is also long and shaped — steering across the whole trajectory —
not a bright-line refusal on any one query.&lt;&#x2F;p&gt;
&lt;p&gt;And we already have the proof of concept: the physical regime does exactly this.
Operative pressure is continuous friction, screening is repeated checkpoints,
licensing is a shaping requirement. None of those is a single “no.” They’re
gentler, distributed management of a threat that unfolds over time.&lt;&#x2F;p&gt;
&lt;p&gt;There’s an objection worth meeting, and it comes from the side I’m defending:
case-by-case consequence estimation makes refusals unpredictable, and the working
scientist may genuinely prefer a bright-line rule to a model forming private
theories about what they intend to do with an answer. A tool that refuses legibly
is easier to build on than one that refuses wisely but inconsistently.&lt;&#x2F;p&gt;
&lt;p&gt;Denmark has gone further than most jurisdictions here. Its biosecurity legislation reaches beyond controlled agents to certain categories of knowledge and skill, obliging a researcher who develops dual-use techniques to approach the authority and establish whether a licence is needed. That’s a capability-based hook rather than an agent schedule, and it’s better law than most countries have. It still doesn’t regulate intent, and can’t.&lt;&#x2F;p&gt;
&lt;p&gt;None of this argues for locking the tools away. It argues for putting the
evaluation effort in the right direction, rather than where our instruments
happen to work. Spend less of it on what a novice can pull out of a chatbot.
Spend more on what a trained professional gets accelerated toward, on whether
our defences stay readable to the systems now designing around them, and — the
part I’d go out on a limb to say almost nobody is talking about in these
discussions — on who, five years from now, will have the training and nothing
left to lose, and how social outcomes will determine security outcomes.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;&lt;em&gt;Written in response to&lt;&#x2F;em&gt; Det biologiske trusselsbillede 2026, &lt;em&gt;Center for Biosikring og Bioberedskab, Statens Serum Institut, June 2026. Their report is available at &lt;a href=&quot;https:&#x2F;&#x2F;www.biosikring.dk&quot;&gt;biosikring.dk&lt;&#x2F;a&gt;. Errors of interpretation are mine.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;https:&#x2F;&#x2F;point.free&#x2F;blog&#x2F;the-ai-bioweapon-risk-isnt-jailbreaks&#x2F;denmark.webp&quot; alt=&quot;&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
</content>
        
    </entry>
</feed>
