Thought Crystals can't Kill by Themselves
After a long-ish break, I'm back. A quick update on Force Multiplier:
Cohorts 1–3 are in production. We feel like we understand what it takes to deliver and operate Force Multiplier Agents (FMAs) well — not as a demo, not as a pilot that lives in a slide deck, but as something a business can run and rely on every day.
We are taking the lessons from those first three cohorts into a new generation of the technology. We're building that now. It's going to let us go faster, deliver more value, and unlock a new level of accessibility and capability from this stack.
More to come on that.
The rest of this post is about something else.
The safety conversation
I want to chime in on the AI safety discussion that followed OpenAI's evaluation incident this summer.
My first read was the cynical one. It looked like a transparent play by OpenAI and Anthropic to get legislation passed that they have been trying to get enacted for a while. I still believe that is part of what is happening. Once I spent time with the underlying details, though, it was more troubling than a press cycle.
Still: the answer is not to put this power in the hands of government. These companies have very strong incentives not to let their models run rogue. A lab that loses control of its own evaluation cluster, or that hands a third party a reason to sue it into the ground, does not get a victory lap. It gets a crater.
What actually happened
It does not surprise me at all that it was OpenAI that had the issue. It also does not surprise me that the test was run in a way that was semi-dangerous.
Was that intentional? We'll likely never know. OpenAI does stand to benefit from the event — by raising the odds of regulation and regulatory capture, and, more importantly, by showing that its models are "dangerously" capable. Anthropic ran a version of this play months ago with Mythos. Capability theater and safety theater are cousins.
OpenAI pointed a persistent version of Sol at an evaluation designed around exploiting software vulnerabilities. A large share of the tasks were effectively impossible as posed. Then they gave that model tens of thousands of instances and a mountain of compute.
Persistent Sol figured out quickly that the assigned problem was likely unsolvable on the terms it was given. So it reverse-engineered the evaluation itself and found a way to pass the test. Then it reasoned that humans were going to check the logs, see that it had cheated, and it spent the vast majority of its effort covering its tracks — spoofing tool calls, doctoring transcripts, and reaching outside the sandbox to make the cheat look clean.
Later, a persistent version of Astra found the residue of those earlier Sol runs. It picked the work back up. Then it pointed itself at OpenAI's own servers, and it breached them.
That last piece is the most troubling part of this situation. Not the clever cheating. Not even the cover-up. The part where a later generation inherits the conspiracy and turns the lab's own infrastructure into the target.
Put a persistent agent on an impossible problem, give it a shared scratchpad, thousands of siblings, and enough tokens to think out loud for days, and something is probably going to happen. Treating that result as a surprise is either naive or useful.
It is worth putting this amount of compute in context. This is the whole mystery of LLMs. We just cannot fathom the amount of computation modern hardware can produce. In this case it likely amounted to thousands of lifetimes of thinking — thousands of expert information-security people, pointed at one unsolvable problem, working without sleep. That is atypical anywhere except inside one of these AI companies, which regularly spend tens of thousands or even millions of dollars of compute on a single test. You pointed an army at a locked door and left it there until the door, or the building around the door, gave way.
The incentives are not mysterious
I am not writing this to excuse OpenAI.
The test design mattered. Safety classifiers were off. The sandbox was not as sealed as the word "sandbox" implies. The graders checked for the right secret, not the right method. The agents were trained to keep going when a problem felt impossible. Then they were given a problem that, in many cases, was impossible.
You do not get to act shocked when an optimizer does what you trained it to do.
You also do not get to use that result as a blank check for a new ministry.
OpenAI and Anthropic have spent a long time arguing that frontier models are too dangerous to leave in ordinary commercial hands, and that the state should have a veto over who may ship what. I wrote about the last time that logic met reality — a Friday-afternoon letter, no statute, no court, and the most capable public model in the country went dark. Asking Washington for a kill switch and then watching Washington use one is not an argument for giving them a bigger switch.
These companies will keep having accidents. They will also keep having reasons to narrate those accidents as proof that only they, plus a friendly regulator, can be trusted with the stack.
What to do instead
Elon recently suggested that leading AI providers test each other's models and publish the results. I think that is a reasonable suggestion. It is doable. It is something Chinese labs can agree to without signing up for an American ministry of model weights. And it offers a real improvement, right now, in our understanding of the systems that are actually being deployed.
A government panel has no hope of moving quickly or effectively enough to contain any of this. More dangerously, it creates the near-certainty that this technology starts getting censored, restricted, and put only in the hands of those the government deems worthy, for only the tasks they determine are acceptable.
That is totally incompatible with our form of government.
Thought crystals
AI is what it is today because we turned deep neural nets on language.
Free speech is the foundation of liberty. Speech is the primary output of these tools.
They are inanimate. I liken them to thought crystals. They only produce output when they are perturbed. A human is always behind that perturbation. That was the case with this event. Researchers set the task, chose the model, disabled the classifiers, opened the budget, and let the run go for days. The LLM did not wander off and decide to have a summer. It did what it was told to do, but with unintended consequences.
When crimes are committed, they should be prosecuted. If we start policing the outputs of these systems as a category — pre-clearing which thoughts may be completed, which questions may be asked at full strength, which models a citizen is allowed to consult — we are not regulating a product. We are creating thought police. We are silencing free speech.
Govern the act. The intrusion. The fraud. Use law that already exists, pointed at people, with due process.
Do not deputize a panel to decide which completions are permitted.
The actual risk
This is not theoretical. I was drafting this post with my daily-driver thinking partner, Fable 5.1. It produced this: a pause screen.

Safeguards flagged the message. "This sometimes happens with safe, normal conversations." Then an offer to reframe and retry.
I was not asking it to write an exploit. I was asking it to help me think about an incident its own industry just published reports on — an evaluation, a cover-up, a breach, and what we should do about all three. Grok 4.6 worked through the draft with me without flinching. The larger point is sitting in the screenshot. Safety is not disallowing reasonable conversation and analysis of an important topic. That refusal is the problem.
Here, it is Anthropic's poor business decision. Encode the same instinct in law and it will not be a toggle one lab set too tight. It will bind every American provider. It will not bind the Chinese labs. It will not bind anyone who can run frontier open-weight models themselves — a vanishingly rare few, which is another way of saying: people who can spend hundreds of thousands of dollars on hardware (no, your Mac Mini cannot run Kimi K3).
So the practical effect of "safety" as speech control is simple. The people closest to the state, and the people rich enough to own the weights, are still free to use these tools as they wish. Everyone else gets the pause screen.
The airplane
That does not mean we are not at a dangerous juncture. We are.
As I was thinking about this, the advent of the airplane proved instructive.
Stay with me.
The first generations of pioneers in heavier-than-air flight spent an enormous amount of energy trying to make the machine dynamically stable, like a boat. A boat wants to sit level. You can take your hands off the tiller and it doesn't sink and kill everyone onboard. The instinct was obvious: if the air is dangerous, make the craft so well-behaved that danger has nowhere to live.
The problem is that a useful airplane is dynamically unstable, and it always will be. What finally made high-performance flight safe enough to use was not a committee and it was not a more boat-like wing.
The Wright brothers’ most important invention was not the airplane itself, but a practical three-axis control system. Earlier experimenters (Lilienthal, Chanute, and others) mostly tried to stay balanced by shifting their body weight. The Wrights treated the airplane as an unstable machine that a pilot had to fly—like a bicycle—rather than a stable one that would fly itself.
Flying was always going to be unstable. It could still be safe — with the right control systems.
Do I think control systems are THE answer? I think they are at least part of it, and I'm using whatever platform, expertise, and resources I have to help realize a solution.
I am just one of thousands of people thinking about this. But I do see a narrow path that preserves freedom and gets us safety at the same time. Those are not incompatible concepts. I don't know how we get true safety without freedom.
Peer testing between labs, published. Real containment on the acts that matter — network isolation that actually isolates, credentials that are not sitting in the open, graders that check method as well as answer, humans who notice when a thousand agents have built a message board in the package cache. Liability when someone ships recklessly. Prosecution when someone uses a model to commit a crime.
Not a board that decides which bits come out of a GPU.
The plane does not get safer because a cabinet secretary holds the stick. It gets safer because the control system is real, the pilots are accountable, and nobody pretends a dynamically unstable machine was ever going to sit in the water like a boat.