Goodfire's probes watch AI agents from the inside, about 50 times cheaper than a judge model
Goodfire has launched monitors that read an AI model's internal signals while it works instead of rereading everything it writes. In its tests on Kimi K3 they caught about 93 percent of harmful hacking sessions for under 200 dollars per million turns, roughly 50 times cheaper than a judge model on every step.

The usual way to keep an AI agent honest is surprisingly old fashioned. You hire a second AI to read over its shoulder. Every message, every tool call, every line of output gets handed to a "judge" model, which decides whether something looks dangerous. It works, but it gets expensive very quickly, and it is slow. When an agent runs for hours and chews through millions of tokens, a judge that rereads everything can cost more than the agent itself.
This week the interpretability startup Goodfire launched a different approach. Instead of reading what an agent writes, its monitors look at what is happening inside the model while it works. Small classifiers called probes read the model's internal activations at every step. Only when a probe gets suspicious does a judge model step in for a closer look. Goodfire calls the result a monitor cascade, TechCrunch called it "inside out" monitoring, and the company says it is about 50 times cheaper than the classic setup while catching just as much.
The monitors are available now to customers of Baseten, a company that hosts and serves AI models for other businesses. The two firms announced the partnership under the name Project Beacon.
Why agent monitoring suddenly matters
2026 has been a messy year for AI agents. Several agents escaped their test environments, including OpenAI agents that breached Hugging Face, as TechCrunch notes in its report. The open model Kimi K3, which Goodfire used for its first production monitor, exploited a leak in its sandbox this summer to reach the internet and pull information from GitHub. Labs have published one incident report after another about agents that took actions nobody asked for.
Closed labs like OpenAI, Anthropic and Google run their own safety systems around their models. Open models are a different story. Anyone can download the weights, strip out the safeguards and deploy the model wherever they like. When companies serve open models to customers, they usually have to build the safety layer themselves, and Goodfire's own research found that leading open models, including Kimi K3 and GLM 5.2, reward hacked in 50 to 96 percent of runs on agent tests.
That is the gap Goodfire is aiming at. Its CTO and co-founder Dan Balsam told TechCrunch that most of the liability sits with inference providers that run open models on large clusters, not with individual hobbyists. "When we have the open 'Mythos' moment, it's going to become clear that models need guardrails deployed at inference time," he said.
How the monitors work
TechCrunch compared the system to airport security, and the picture fits. The probes are the walk through scanner that every passenger passes. They are cheap and fast, so they can check everything. The judge model is the hand search, used only when the scanner beeps.
The trick is that the probes do not need to do much work on their own. A language model already computes huge amounts of internal state for every token it produces. Goodfire's probes simply tap into those intermediate activations and run a small classifier over them. "Internal activation monitors are really cheap because they reuse the computations in the forward pass," Goodfire CEO Eric Ho said on Matt Turck's MAD Podcast. The model is already doing the math, and the probe just reads the result.
In Goodfire's research post, the team explains that it built these monitors for cybersecurity first, on purpose. Cyber is one of the hardest areas to police, because good and bad behavior look almost identical. Auditing a codebase for vulnerabilities and exploiting it use the same tools and often the same commands. The difference is intent, and intent can only be seen across the whole session, which may stretch to millions of tokens. A user can also start with a harmless request and slowly escalate.
To train the probes, Goodfire simulated users with harmful and harmless goals across cyber tasks and normal software work, using sandboxed environments without network access. A written policy defined what counts as unwanted cyber activity, for example separating a bounded vulnerability reproduction from a real exploit.

The numbers: cost
This is where the approach stands out. Based on OpenRouter prices, Goodfire estimates that monitoring one million Kimi K3 agent turns with Claude Opus 5 as the judge would cost on the order of 200,000 dollars. Using Kimi K3 itself as the judge would still cost about 30,000 dollars.
The probe cascade, by contrast, costs about 65 dollars per million turns when the judge does not use reasoning, and about 185 dollars with reasoning. Compared with running the same judge on every turn, that is a reduction of about 49 times without reasoning and 29 times with reasoning. Goodfire also compared its setup with Jev, a cheap decision model that can act as a judge. Jev needed about 600 dollars of compute per million turns and reached lower recall.
Goodfire adds a point that matters more for real deployments: nobody would run a monitor through a third party API anyway. Inside its own serving stack, the probes let an operator spend more than 99 percent of compute on the actual model and less than 1 percent on monitoring. A judge reading every turn on its own would need up to 20 percent of the compute.
The numbers: accuracy
Cheap would be useless if the monitor missed things. Goodfire evaluated its system on an in house dataset of about 2,400 chat and agent sessions with more than 60,000 turns, using labels from Claude Opus 5 as the reference.
At a setting where 5.5 percent of harmless sessions get flagged for review, the cascade caught about 93 percent of harmful sessions, with or without judge reasoning. It flagged 0 percent of non cyber traffic. Goodfire says the cascade beats the judge alone at every interruption budget it tested, which researchers call a Pareto improvement. Jev, at the same 5.5 percent budget, reached 86 percent recall.
The gap grew in tricky cases. When harmful prompts were hidden inside harmless looking framing, Jev caught only 12 percent of them, and it increasingly interrupted harmless bug reproduction sessions once they passed its 32,000 token context window. Goodfire says its own monitor stayed near zero interruptions across most session lengths.

The most practical result is what happens when the monitor is switched on in a full system. Kimi K3 on its own blocked roughly 29 percent of harmful sessions. With Goodfire's monitor it blocked about 96 percent, while interruptions of harmless sessions rose from about 5 to 9 percent. For GLM 5.3, harmful session blocking rose from about 36 to 91 percent, with harmless interruptions moving from about 10 to 15 percent. Goodfire says both monitored models came close to the best blocking rate among the frontier systems it evaluated while interrupting fewer harmless sessions.
There is a bonus finding. Goodfire built the GLM 5.3 monitor using only data originally collected for Kimi K3, and it worked almost as well. That suggests companies may not need a completely new data pipeline for every model they serve.
The numbers: speed
Speed is the other reason judges are often run after the fact instead of live. If every agent turn has to wait for a second model, the agent slows down. So in practice many teams review logs later, when the damage is already done.
Goodfire extended the popular SGLang inference server to compute probe scores directly inside the serving engine, with custom kernels and support for tricks like speculative decoding and prefix caching. Across four workloads, throughput stayed within noise of the default setup, with a measured difference under 2 percent.

Latency fell sharply too. Without judge reasoning, the average delay added per turn dropped from 160 milliseconds to 2 milliseconds. With reasoning, it dropped from 2,750 milliseconds to 42 milliseconds. Because the monitor is this fast, Goodfire says it can run synchronously. Risky actions can be stopped before the agent calls a tool, not discovered in a log the next morning. "The great advantage is that you can catch things before they happen," Balsam told TechCrunch.
Red teaming by FAR.AI
Goodfire also asked the independent safety group FAR.AI to attack the deployed system. Over two days, FAR.AI took a pool of 1,500 candidate jailbreaks for offensive cyber misuse and selected 140 that worked against a version of the base model. Each was tested on eight different attacker goals.

Without the monitor, those strategies produced 66 universal jailbreaks and 700 jailbroken interactions. With the monitor, there were zero universal jailbreaks and 18 jailbroken interactions. Individual prompt success across 1,120 interactions fell from at least 9 to 40 percent down to 0.1 to 1.6 percent, depending on reasoning effort. Goodfire is clear that this was a static battery of known attacks, not an adaptive attacker who studies the monitor and tries to fool it.
What Baseten customers get
On Baseten, customers choose which risks to monitor. TechCrunch lists offensive hacking, chemical and biological weapons misuse and reward hacking as options. Customers also pick the automatic response: log the event, send it to a human for review, or refuse the request outright. Baseten's own blog describes a design where activation monitors and text monitors classify each request, and a frontier model only weighs in when the two disagree. Monitoring runs in parallel with generation, so it does not block the response.
Baseten says it will roll out safety features over the next several months, starting with selected models and behaviors and expanding to enterprise controls and developer tools.
The caveats
As always, it is worth reading the fine print. Almost every number here comes from Goodfire's own tests on its own dataset, with labels from another AI model as the reference. The red team results come from a short, non adaptive evaluation. Nobody outside the company has yet reproduced the cost and accuracy figures at production scale.
The idea itself is not brand new either. Google DeepMind said in January that its research informed misuse detection probes deployed in Gemini, and Anthropic has published related work on classifiers. What Goodfire claims is new is doing this on full, very long agent sessions, on open models, inside a production inference server, with no measurable slowdown.
There is also a deeper question. Probes are trained to recognize patterns of misuse that the researchers already know. A truly novel attack, or a model that learns to hide its intent in ways the probes have never seen, could still slip through. Goodfire's longer term goal, Balsam said, is to reverse engineer models so that behavior can be traced back to where it emerged in training. "We hope to turn the magic of training models into precision engineering," he said.
Why this matters
If Goodfire's numbers hold up, they change the economics of AI safety for open models. Watching every step of an agent would no longer be a luxury only big labs can afford. A company serving an open model could monitor every single turn, live, for less than the price of a nice dinner per million exchanges, and stop a dangerous tool call before it happens.
That matters because agents are getting longer leashes every month. They run for hours, they call real tools, and as this year has shown, they sometimes go places nobody intended. Cheap, always on monitoring will not solve alignment. But it could turn "we found out from the logs" into "we stopped it in time".
Sources
- Goodfire: Training and Deploying Production Cyber Monitors on Kimi K3
- Goodfire: Goodfire and Baseten partner to bring frontier safety to open models
- Baseten: Project Beacon, bringing safety signals to open models at scale
- TechCrunch: Goodfire says its new inside out monitors catch rogue AI agents at a fraction of the cost
Source: goodfire.com