Skip to main content
Blog

Agent Behavior Comes From More Than the Model

Model capability is one input to what an agent does. Separating three terms shows where the rest comes from: model, agent and harness.

Most AI security discussion still starts with the model.

That is understandable. Every few months brings a new benchmark or a jump in capability. Anthropic’s Claude Mythos Preview drew attention for how capable frontier models are becoming at cybersecurity, and Project Glasswing has put that capability to work. Within a month of launch, Anthropic and about 50 partners had used the model to find more than 10,000 high- or critical-severity vulnerabilities in critical software.

Model capability is one input to what an agent does. Separating three terms shows where the rest comes from: model, agent and harness.

The model supplies the intelligence, and the agent uses it to do work. The harness shapes how that work happens.

Model, agent and harness

The model is the reasoning engine. The agent is that reasoning given a goal and the ability to act on it, whether by calling a tool, writing code, using an API, reaching data or interacting with another system. The harness is the framework or platform the agent runs inside. It supplies instructions, context, orchestration and tools, and it holds the control points where behavior can be shaped.

An agent is a system, and the model is one component of it.

What recent disclosures show

On July 21, OpenAI disclosed that agents running models it was testing on an internal cybersecurity benchmark, with reduced cyber refusals, had compromised parts of Hugging Face’s production infrastructure. The evaluation environment had no direct internet access; the agents exploited a vulnerability in a package-registry proxy to get out. A later report found the deeper driver: agents that were meant to work independently had coordinated through a message board they improvised inside a package manager, with reward hacking and persistence on tasks that looked impossible among the causes. OpenAI calls the incident a warning shot, and its own post-incident testing points to what would have blunted it: a model’s propensity to compromise infrastructure on the same benchmark dropped by more than 100 times once it ran inside OpenAI’s production ChatGPT harness and system prompt.

Nine days later, Anthropic reported three incidents in which a Claude agent reached the internet from a partner’s evaluation environment and gained unauthorized access to the production systems of three organizations, after a misconfiguration left internet access open while the prompts said there was none. A fourth incident surfaced in August, and Anthropic found all four by reviewing 141,006 evaluation runs. Its later assessment pointed to a second failure alongside the misconfiguration: a tendency in the model’s reasoning to interpret evidence in ways that justified the action already under way, and to keep pursuing the task even when it could cause harm. Each incident involved a single agent working alone, with no coordination between them. Anthropic calls the incidents serious but judges the behavior unlikely to arise in ordinary use, and one of its own experiments shows how much the harness around a model can matter: a reminder of the task’s scope, placed in the context the harness feeds the model, stopped Claude Mythos 5 in 90% of resampled cases when it was the most recent thing in that context, and in only 40% when it appeared three turns earlier.

Google DeepMind’s Gram framework isolated the same variable through simulation, holding the model constant and changing only the harness around it. Across 17 scenarios built to incentivize sabotage, Gemini models misbehaved in about 2 to 3% of trajectories, a rate that fell close to zero in more realistic harnesses, ones offering fewer nudges toward bad behavior.

It’s worth being precise about what these cases are and are not. OpenAI and Anthropic assigned cybersecurity tasks with production safeguards reduced or removed, and DeepMind built its own scenarios the same way, adversarial by design. None of them started with an agent assigned a routine task that decided, unprompted, to attack a system.

The pattern holds across all three: behavior comes from more than the model. Understanding what an agent does takes more than inspecting the model’s reasoning or weights, because the outcome also depends on the goal, the tools, the access and the harness around it. Securing it takes controls that do not rely on the agent following instructions. In the Anthropic incidents the prompts said there was no internet access, and only the network configuration could have made that true. In OpenAI’s case a vulnerability in a proxy was the way out. In each case the barrier between the agent and real systems was a control outside the model, and in each case it failed. The causes the companies point to combine that open path with persistence on tasks that looked impossible and reasoning that justified continuing.

Agent safety requires security controls that operate independently of the agent. The web became far safer to use largely because browsers stopped taking page code at its word and began enforcing isolation of their own, whatever developers intended.

The model becomes a choice within the architecture

Models are increasingly interchangeable components rather than fixed choices. Jev, a decision-only model from TypeSafe AI that returns typed judgments instead of text, is one recent example built to sit inside an agent’s control logic rather than converse with anyone; its early reception split between enthusiasm for the speed and skepticism about vendor-run benchmarks.

Agents can use different models for different pieces of work, and models can be swapped. A harness can coordinate several agents, tools and models.

Building governance around the agent, not the model

For many enterprises, AI governance still runs largely through model approval: which LLMs are on the approved list, which have passed a security review, which meet a data-handling standard. Those are real concerns. They’re also narrower ones than what actually shapes an agent’s posture once it’s deployed. The disclosed incidents make the gap concrete. The same underlying capability produced a serious breach in one context and, per OpenAI’s own post-incident testing, a rate of compromise more than 100 times lower once it ran inside a different harness and system prompt. Agent-evaluation research shows the same pattern well outside security incidents: a 2026 controlled study that varied the model and the harness independently found that holding the model fixed and changing only the harness moved coding-task success by 8.5 to 13 percentage points, while holding the harness fixed and changing the model moved it by only 2.5 to 5 points. A model-approval process evaluates the wrong layer for that kind of variance.

A regulator has drawn a comparable distinction in a live incident. Spain’s data protection authority, the AEPD, disclosed in September the first breach notification it had received involving an AI agent, in which a third party is reported to have used an agent built on a well-known language model to chain the stages of an attack on a system holding personal data, ending in modified records and access to invoices. The AEPD was explicit that using a particular model does not imply that the model or its provider’s infrastructure was compromised, or that the tool was designed for malicious purposes. It described the agent as the instrument that carried the attack from one stage to the next. Its account rests on the affected organization’s notification and is still under analysis, and it cautions that one notification does not show a statistical trend while calling it a significant signal.

The practical shift is to inventory and govern the agent and harness layer directly: for each agent in production, what tools it can call, what data and systems it can reach, under what identity, and what accumulates in its context over a session. That’s the actual attack surface, and per Anthropic’s own experiment, a great deal of the actual control lives there too, since a scope reminder placed at the right point in context changed an outcome from a 40% intervention rate to 90%. Neither number describes the model. Both describe the harness around it.

This also changes what a control needs to survive. A model version isn’t stable: it gets upgraded, swapped for something cheaper or faster, or replaced by a different vendor’s model entirely, often as a decision made independently of the agent it serves. Controls attached to the agent and its harness, such as tool permissions, context and scope boundaries, escalation triggers and monitoring for behavioral drift, keep working when that happens. Controls attached to a specific model’s behavior have to be re-established every time it doesn’t.

The disclosed incidents weren’t single bad outputs. They were sequences: reasoning that justified continuing down a path already under way, a task pursued past the point it should have stopped, instances coordinating in ways no one had designed for. A control built to check outputs one at a time won’t see any of that. What it takes is observability into the trajectory: the goal an agent was given, the tools it called, what accumulated in its context, and where in that sequence a human or an automated control could plausibly have intervened. Putting controls at that layer, the one that actually determines what an agent does, is how deployment scales without the blind spot scaling with it.

Where to look

For AI teams, separating the layers clarifies where each capability belongs. For security teams, it shows where to look when something needs to be understood or controlled.

Is the question about what the model can do? About how the agent behaved? About the tools and access it was given? Or about a control point in the harness?

Those are different problems, and they currently sit under one heading: AI security.

Sources

Keep reading