Ledger A · The security record
Which agents exist, what they can reach, what they did, how risky it was.
- Owned by security
- Reported to the board as posture
- Silent on cost
What we’re solving
The business shipped agents in months. Making sense of the roll out has become the most important job in the business.
What changed with agents
Agents research, analyze, build and communicate. They make decisions, call tools and take actions across the same systems, data and workflows people depend on every day. This is not a copilot in a sidebar. It is work moving through software that decides for itself what to do next.
Which means every organization now needs a new understanding of the same operating environment: which agents are working, what they are doing, what they can access, how they are behaving, what is happening as they work, what it costs, and where attention needs to go.
Those questions span security, AI, technology and the business itself. Almost no organization has governance infrastructure built to answer them, because agents are systems, not surfaces – their configuration, identity, tools, permissions and activity span cloud, code, endpoint and browser at the same time.
01
Most teams cannot compile a complete, current list of what agents are deployed, where they operate and who owns them — so posture, compliance and control decisions are all made on incomplete data.
327% more agents than one customer’s own security team estimated
02
Skills, extensions, plugins, packages, SaaS connectors, direct API calls. Gateways see one of those. Agents retain access inconsistently across all of them.
03
Endpoint-only has a ceiling. Gateways have a lane. Neither can govern a system that lives everywhere at once.
04
Developer agents wasting tokens on side projects. Agents that keep running long after they should have stopped. Three departments paying separately for the same work. It comes up in every other first meeting, and often before security does.
What the inventory shows
Every organization we have deployed into found more agents than it expected. Not marginally more. The gap between the inventory and the estate is where risk, spend and accountability all go missing at once.
In the register reviewed, owned, approved
Running anyway shadow, citizen-built, developer preference, duplicated across departments
Unapproved customer-facing agents processing sensitive complaints. Developer agents running side projects on company tokens. A marketing agent built by someone who never knew a review process existed. None of it appears in the register, and none of it stops running while the register is being written.
Without behavioral context, all three rows are the same row.
A coding agent interacting with source code while holding credentials associated with a competitor reads as normal development activity. An agent pushing corporate secrets to a public repository reads as a routine commit. The difference is context, and context is exactly what a boundary log strips out.
Then there is the record itself: not every agent produces logs, and those that do often permit the session to be rewritten. Without an immutable behavioral record, incident response begins after the damage rather than during it.
Agents are systems, not surfaces
| Tool | Cloud | Code | Endpoint | Browser |
|---|---|---|---|---|
| Total agent coverage | covered | covered | covered | covered |
| Endpoint sensor | not covered | not covered | covered | not covered |
| MCP gateway | covered | not covered | not covered | not covered |
| Cloud posture tool | covered | not covered | not covered | not covered |
| Identity platform | permissions only, not activity | permissions only, not activity | permissions only, not activity | permissions only, not activity |
| Cost dashboard | not covered | not covered | not covered | not covered |
Every one of these tools is doing its job correctly. The problem is that the agent is a single system and each tool is looking at one wall of it. Assembling five partial answers by hand, weekly, is not governance – it is archaeology.
Agent spend is now a board-level line item, and it is the one line the person accountable for it cannot break down. Not because the number is wrong - because the number has no work attached to it.
What you get today
API and model usage £418,900
One aggregate. No agent. No owner. No task. No answer to “was any of that worth it?”
What the question actually needs
| Agent | Work it was doing | Owner | Spend |
|---|---|---|---|
| claims-triage | Claims intake | Claims ops | £96,400 |
| market-research | Competitor scan | Strategy | £61,200 |
| pr-review-agent | Retrying a failed job, 3,400 times | Platform eng | £73,800 |
| unnamed-dev-agent | Personal side project | Unattributed | £44,100 |
| 3 × content-drafter | The same task, three departments | Marketing, Sales, CS | £143,400 |
Three of these five rows are waste, and none of them are visible until spend is attached to behavior.
Mode 01
Developer agents running side projects on company tokens. Perfectly reasonable behavior by the person, entirely invisible to the budget holder.
Mode 02
Uncontrolled cost scaling from agents that kept running when they should have exited – retry loops, orphaned child sessions, work that finished hours ago.
Mode 03
Duplicate agents doing identical work in different departments, because no department can see the others’ estate.
Mode 04
Unnecessary frontier models for routine tasks. Employees forgetting to reset after a complex and intensive project.
The structural problem
Security tools do not surface financial risk.
Cost tools do not surface security risk.
The same agent, doing the same thing, appears in both records – and the two records never meet.
Ledger A · The security record
Which agents exist, what they can reach, what they did, how risky it was.
Ledger B · The financial record
How many tokens, on which model, in which month, at what total.
Not hypothetical
Every one of these is drawn from a real deployment. Every one of them looked like ordinary activity right up to the moment somebody understood the context.
Critical Supply chain
It looked legitimate until somebody understood what the agent was doing with it. Skills sit in the agent’s own configuration and generate no outbound traffic for a gateway to intercept.
High Data exposure
Customer-facing, never reviewed, handling exactly the material a review would have caught. It was not in any inventory.
Critical IP leakage
In the commit history it reads as routine work. The distinction only exists if you can see what the agent was holding at the time.
Medium Propagation
Agent-to-agent spread through the same mechanisms that make agents useful. Nothing in the traditional stack is watching that path.
High Insider threat
An employee, an unlicensed agent, an account tied to a rival. Customer data and company IP both exposed. Endpoint and network controls saw normal development.
Medium Financial risk
No oversight on cost, no attribution, no cap. This one comes up in every other first customer meeting — and it is the one nobody classifies as a security problem until they see whose credentials paid for it.
A patchwork approach offers no audit trail, no behavioral history, no cost attribution and no fast path to remediation for any of the six.
Bottom line:
Outcome one
The bottleneck. Adoption slows, the team becomes the thing standing in the way of the strategy the board announced.
Outcome two
Unmanaged expansion. The estate grows without review, and the first accurate inventory is produced by an incident.
Both outcomes create risk. The choice is only forced because the evidence needed to make a real decision does not exist yet.
Quarter 1
This is the cheapest moment to establish governance, and the moment it feels least necessary.
Quarter 2
Business teams build their own. Developers adopt the tool they prefer over the one they were given. Nobody is doing anything wrong.
Quarter 3
Finance asks for a breakdown. The honest answer is a total and a shrug. This is usually the quarter the cost conversation opens the security conversation.
Quarter 4
Behavioral history cannot be reconstructed retrospectively. Whatever was not recorded at the time is simply gone.
33%
of enterprise software applications will include agentic AI by 2028
Gartner
15%
of day-to-day operations completed autonomously by 2028
Gartner
40%
of agentic AI projects will be cancelled due to unmanaged risk
Gartner
12
months in which every major platform shipped agent capability
Anthropic, Microsoft, OpenAI, GitHub, AWS
Every requirement on this page points the same direction: understanding has to come from inside the agent’s own execution path, not from a boundary near it.