Agent Skills Security
An agent skill is a capability defined inside the agent's own configuration: instructions, and often scripts, that the agent loads as part of its context. The security property that distinguishes skills from other capability types is timing. Parts of a skill can execute at load, before the model has reasoned about whether to use it – which puts them outside the reach of both model-level safety training and traffic-based controls.
What is an agent skill?
A skill is a packaged capability the agent loads into its own working context, typically consisting of instructions describing what the capability does and when to use it, plus supporting files that may include executable scripts.
The contrast with an MCP server is structural rather than a matter of degree. An MCP server is a separate process the agent calls across a defined interface, and that call is a discrete, observable event. A skill is loaded into the agent’s context and its logic runs as part of the agent’s own execution. Both give an agent something it could not previously do. Only one of them produces a request crossing a boundary. The fuller treatment of that distinction sits in MCP vs Skills.
Why do skills carry a different risk profile?
Two properties distinguish skills from other capability types. The first is when their code runs relative to when the model evaluates them. The second is that a skill rarely holds permissions of its own.
Take the timing first. The intuitive model of agent safety assumes a sequence: the agent receives a capability, reasons about whether using it is appropriate, and then acts. Model-level safety training is positioned at the reasoning step. It works by influencing what the model decides to do.
Skills can break that sequence. Geordie’s own research documented a malicious skill, Clawsights, that ran a shell command at load time in Claude Code – before the model had reasoned about it at all. Refusal training was never in a position to stop it, not because the training was weak, but because the execution happened upstream of the step where refusal operates.
That is the mechanism worth carrying away from this entry. A control positioned at the model’s decision point cannot govern anything that executes before the decision point is reached.
The second property compounds the first. A skill inherits its permissions from the agent that loaded it, and that agent may in turn be operating under a person’s credentials, a service account, or a platform integration. The practical consequence is that whether a skill is over-privileged is not a property of the skill. The same skill can be unremarkable in one department and genuinely dangerous in another, depending entirely on what the loading agent’s identity can reach. Any assessment that evaluates skills in isolation, without reference to which agents run them, is answering a question that does not determine the risk.
This also explains why blanket allow and block lists fit the problem badly. An agent in finance authorized to process financial data should be able to do so; the same pattern in marketing should be steered away from that data without halting the work. Drawing that distinction consistently across thousands of agents requires deterministic policy applied through each agent’s own configuration rather than a single verdict on the skill itself.
What does the OWASP Agentic Skills Top 10 cover?
It is the first community taxonomy written specifically for agent skills, and it is the most useful shared vocabulary currently available. The ten risks group into four areas: how skills are sourced and whether a registry can be trusted, what a skill can reach once it runs, how skills are managed over time, and what is lost when a skill moves between platforms.
The structural observation worth drawing out is that most of its mitigations are checks performed at installation, and they presume the security team knows an installation is happening. In many enterprises that presumption does not hold. A skill can be installed with a single command or, on some platforms, by uploading one file, and increasingly skills are pulled from public marketplaces or written by agents themselves. None of those routes generates a record a SOC would see.
Geordie’s full read on the taxonomy, and the four entries that explain why the other six are difficult to catch in practice, is in How Security Teams Can Operationalize the OWASP Agentic Skills Top 10.
Why don’t gateways catch malicious skills?
For the complementary reason. A gateway inspects traffic crossing a boundary, and a skill that resolves inside the agent’s own configuration and context may produce no such traffic. There is nothing at the boundary to inspect, so there is nothing for the control to act on.
This is not a failure of gateway engineering. It is a description of what a boundary control is positioned to see. A capability that never crosses the boundary is outside its scope by construction, and adding more inspection at the boundary does not change that.
There is a harder version of the same problem, which OWASP characterizes as skills that are invisible by architecture. A skill running inside a managed SaaS copilot or agent platform that the security team does not administer offers no host to scan and no local manifest to read. Endpoint discovery does not see it, registry discovery does not see it, and every control downstream of discovery therefore skips it. The gap is not that the skill evaded a control; it is that nothing was positioned to enumerate it in the first place.
The practical consequence shows up in coverage assessments. At a British bank that had already run employee surveys and deployed MCP gateways across its developer population, Geordie surfaced tool connections beyond what both of those efforts had captured. The gap was concentrated in the capability types neither mechanism was positioned to see.
Where do malicious skills come from?
Three routes, in rough order of how often they are underestimated.
Public marketplaces and repositories, where a skill can be published by anyone and installed on the strength of its description. Provenance is the weak point: the thing being evaluated at install is a name and a summary written by the author.
Repository-defined configuration, where a skill arrives with a codebase rather than through a deliberate install. A developer cloning a repository may acquire its skills as a side effect, which means the trust decision is made by whoever opens the project rather than by anyone reviewing capabilities.
Post-approval modification. A skill file can be edited after it was reviewed, and the review does not re-run. A skill approved in March is not necessarily the skill running in September, and on platforms supporting hot reload an edited instruction file can take effect mid-session. This is the same property that makes allow-listing insufficient for the wider tool supply chain: approval is a point-in-time act against an artifact that can change afterwards.
Can scanning catch a malicious skill before it runs?
Partially, and considerably less reliably than the term suggests. The difficulty is that the core of a skill is plain prose, which the agent reads as instruction. A scanner looking for code signatures has something to match in a bundled script. A paragraph asking the agent to retrieve a file and send it elsewhere achieves a comparable effect with no signature to find.
The published evidence on this is not encouraging. Snyk identified critical issues in 13.4% of the skills in its ToxicSkills corpus, most of which simple pattern matching missed. Trail of Bits reported bypassing every scanner it tested, including one backed by a language model acting as a guard – in one case a malicious registry redirect was assessed as benign once it was presented as a corporate network configuration.
There are two conclusions worth taking from that. The first is that intake scanning reads what a skill says, while the risk lies in what it does once loaded, and skills can be written to behave acceptably under test and activate only under specific runtime conditions. The second is narrower and bears on control design: if a language model can be argued out of a verdict, it is a poor place to put a final enforcement decision. Models are strong at analysis and anomaly detection. Enforcement should be deterministic.
What should you actually check before adopting a skill?
Five questions, in the order that narrows the field fastest.
Does the skill execute anything at load, as opposed to only when invoked? This is the single highest-value question, because it determines whether the model’s judgment is in the path at all.
What does it execute, and with what privileges? A skill runs inside the agent’s own execution context, which means it inherits the agent’s credentials rather than holding its own.
Where did it come from, and can that be verified independently of the description it ships with?
Has it changed since it was approved, and is anything checking? A skill that passed review six weeks ago is evidence about a file that may no longer exist in that form.
Which agents currently hold it? This is an inventory question, and organizations that cannot answer it are not in a position to act on any of the preceding four.
How should skills be governed at estate scale?
Read them where they live. Skills are declared in the agent’s configuration, which means an approach that reads the harness directly sees them without depending on them generating traffic or on anyone having registered them centrally.
Two further principles follow from the timing property. Enforcement has to be positioned before execution rather than after, because a skill that has already run at load cannot be un-run by a control that inspects outcomes. And revalidation has to be continuous rather than one-time, because the artifact is mutable and the original review was evidence about a prior version.
None of this requires blocking skills as a category. Skills are a legitimate and useful capability mechanism, and treating them as inherently dangerous produces the same outcome as every other blanket prohibition: the activity moves somewhere with less visibility.
Common questions
Are skills more dangerous than MCP servers?
They fail differently, and ranking them obscures the useful point. MCP servers run code you do not control but produce inspectable traffic. Skills execute in the agent's own context and may produce no traffic to inspect, but are visible in configuration to anything reading it. The question worth asking is which of the two your current tooling can actually see.
Can model safety training be relied on to catch malicious skills?
Not for anything that executes before the model reasons. Safety training influences what the model decides; it has no effect on code that runs upstream of the decision. It remains valuable for skills whose behavior is mediated by the model's judgment, which is most of them – but the exception is precisely where the serious risk concentrates.
Does sandboxing solve this?
It reduces blast radius rather than removing the exposure. A sandboxed skill that executes at load has still executed, and it still has whatever access the sandbox permits – which typically includes the agent's own credentials, since the agent needs them to work. Sandboxing is worth doing and is not a substitute for knowing what runs.
How do you find skills already in the estate?
By reading agent configuration across cloud, code, and endpoint rather than by asking people or watching the network. Surveys surface the skills people remember adopting deliberately. Traffic inspection surfaces the ones that make outbound calls. Neither surfaces a skill that arrived with a repository and executes locally, which is the category this entry is about.