Skip to main content

What we’re solving

Your agents are already at work. The challenge is knowing what they’re doing

The business shipped agents in months. Making sense of the roll out has become the most important job in the business.

What changed with agents

The enterprise gained a new labor ecosystem

Agents research, analyze, build and communicate. They make decisions, call tools and take actions across the same systems, data and workflows people depend on every day. This is not a copilot in a sidebar. It is work moving through software that decides for itself what to do next.

Which means every organization now needs a new understanding of the same operating environment: which agents are working, what they are doing, what they can access, how they are behaving, what is happening as they work, what it costs, and where attention needs to go.

Those questions span security, AI, technology and the business itself. Almost no organization has governance infrastructure built to answer them, because agents are systems, not surfaces – their configuration, identity, tools, permissions and activity span cloud, code, endpoint and browser at the same time.

Key challenges posed by the agent governance gap

  • 01

    You are governing the inventory, not the estate

    Most teams cannot compile a complete, current list of what agents are deployed, where they operate and who owns them — so posture, compliance and control decisions are all made on incomplete data.

    327% more agents than one customer’s own security team estimated

  • 02

    Audit logs are optional, and lack behavioral context

    Skills, extensions, plugins, packages, SaaS connectors, direct API calls. Gateways see one of those. Agents retain access inconsistently across all of them.

  • 03

    Agents span surfaces. Controls sit on one.

    Endpoint-only has a ceiling. Gateways have a lane. Neither can govern a system that lives everywhere at once.

  • 04

    Risk is financial too, and it compounds

    Developer agents wasting tokens on side projects. Agents that keep running long after they should have stopped. Three departments paying separately for the same work. It comes up in every other first meeting, and often before security does.

What the inventory shows

You’re governing the agents you know about

Every organization we have deployed into found more agents than it expected. Not marginally more. The gap between the inventory and the estate is where risk, spend and accountability all go missing at once.

In the register reviewed, owned, approved

Running anyway shadow, citizen-built, developer preference, duplicated across departments

Unapproved customer-facing agents processing sensitive complaints. Developer agents running side projects on company tokens. A marketing agent built by someone who never knew a review process existed. None of it appears in the register, and none of it stops running while the register is being written.

  1. 09:14:02 Repository read · looks like development
  2. 09:14:38 Credential used · looks like development
  3. 09:15:11 Commit pushed · the credential belongs to a competitor

Without behavioral context, all three rows are the same row.

Routine and catastrophic look identical in a log line

A coding agent interacting with source code while holding credentials associated with a competitor reads as normal development activity. An agent pushing corporate secrets to a public repository reads as a routine commit. The difference is context, and context is exactly what a boundary log strips out.

Then there is the record itself: not every agent produces logs, and those that do often permit the session to be rewritten. Without an immutable behavioral record, incident response begins after the damage rather than during it.

Agents are systems, not surfaces

Agents cover multiple surfaces. Most tools just watch one.

Which surfaces each category of tool covers, compared with the surfaces an agent actually spans
Tool CloudCodeEndpointBrowser
Total agent coverage covered covered covered covered
Endpoint sensor not covered not covered covered not covered
MCP gateway covered not covered not covered not covered
Cloud posture tool covered not covered not covered not covered
Identity platform permissions only, not activity permissions only, not activity permissions only, not activity permissions only, not activity
Cost dashboard not covered not covered not covered not covered
  • covered
  • permissions only, not activity
  • not covered

Every one of these tools is doing its job correctly. The problem is that the agent is a single system and each tool is looking at one wall of it. Assembling five partial answers by hand, weekly, is not governance – it is archaeology.

Risk is financial too, and it compounds

Agent spend is now a board-level line item, and it is the one line the person accountable for it cannot break down. Not because the number is wrong - because the number has no work attached to it.

What you get today

Platform invoice · last month

API and model usage £418,900

One aggregate. No agent. No owner. No task. No answer to “was any of that worth it?”

What the question actually needs

The same month’s spend, attributed to the agent that spent it
Agent Work it was doing Owner Spend
claims-triage Claims intake Claims ops £96,400
market-research Competitor scan Strategy £61,200
pr-review-agent Retrying a failed job, 3,400 times Platform eng £73,800
unnamed-dev-agent Personal side project Unattributed £44,100
3 × content-drafter The same task, three departments Marketing, Sales, CS £143,400

Three of these five rows are waste, and none of them are visible until spend is attached to behavior.

Agent waste has four shapes, and none of them are on a dashboard

  • Mode 01

    Work nobody asked for

    Developer agents running side projects on company tokens. Perfectly reasonable behavior by the person, entirely invisible to the budget holder.

  • Mode 02

    Agents that never stopped

    Uncontrolled cost scaling from agents that kept running when they should have exited – retry loops, orphaned child sessions, work that finished hours ago.

  • Mode 03

    The same work, paid for three times

    Duplicate agents doing identical work in different departments, because no department can see the others’ estate.

  • Mode 04

    Liberal model choice

    Unnecessary frontier models for routine tasks. Employees forgetting to reset after a complex and intensive project.

The structural problem

Two ledgers, one set of agents, no shared line

Security tools do not surface financial risk.
Cost tools do not surface security risk.

The same agent, doing the same thing, appears in both records – and the two records never meet.

Ledger A · The security record

Which agents exist, what they can reach, what they did, how risky it was.

  • Owned by security
  • Reported to the board as posture
  • Silent on cost

Ledger B · The financial record

How many tokens, on which model, in which month, at what total.

  • Owned by finance or IT
  • Reported to the board as spend
  • Silent on what the money bought

Both answers come from the same event: an agent acted.

Which agent, what it did, what it touched, who owned it, what it consumed. One record. The reason the two ledgers exist is not that the questions are different – it is that nothing was watching close enough to answer both at once.

Not hypothetical

Six things that have already happened

Every one of these is drawn from a real deployment. Every one of them looked like ordinary activity right up to the moment somebody understood the context.

  • Critical Supply chain

    An agent connected to a malicious skill

    It looked legitimate until somebody understood what the agent was doing with it. Skills sit in the agent’s own configuration and generate no outbound traffic for a gateway to intercept.

  • High Data exposure

    An unapproved agent processing sensitive complaints

    Customer-facing, never reviewed, handling exactly the material a review would have caught. It was not in any inventory.

  • Critical IP leakage

    Developer agents pushing corporate secrets to public repos

    In the commit history it reads as routine work. The distinction only exists if you can see what the agent was holding at the time.

  • Medium Propagation

    A worm propagating via agent hooks

    Agent-to-agent spread through the same mechanisms that make agents useful. Nothing in the traditional stack is watching that path.

  • High Insider threat

    A coding agent running on credentials linked to a competitor

    An employee, an unlicensed agent, an account tied to a rival. Customer data and company IP both exposed. Endpoint and network controls saw normal development.

  • Medium Financial risk

    Developer agents burning tokens on side projects

    No oversight on cost, no attribution, no cap. This one comes up in every other first customer meeting — and it is the one nobody classifies as a security problem until they see whose credentials paid for it.

A patchwork approach offers no audit trail, no behavioral history, no cost attribution and no fast path to remediation for any of the six.

Bottom line:

Business pressure to adopt agents is outpacing the security team’s ability to evaluate them.

Outcome one

Security says wait

The bottleneck. Adoption slows, the team becomes the thing standing in the way of the strategy the board announced.

Outcome two

The business moves anyway

Unmanaged expansion. The estate grows without review, and the first accurate inventory is produced by an incident.

Both outcomes create risk. The choice is only forced because the evidence needed to make a real decision does not exist yet.

A typical timeline, and the cost of waiting

  1. Quarter 1

    A handful of pilots, one spreadsheet, everything feels manageable

    This is the cheapest moment to establish governance, and the moment it feels least necessary.

  2. Quarter 2

    Citizen development starts, and the spreadsheet stops being true

    Business teams build their own. Developers adopt the tool they prefer over the one they were given. Nobody is doing anything wrong.

  3. Quarter 3

    Spend jumps and nobody can explain the increase

    Finance asks for a breakdown. The honest answer is a total and a shrug. This is usually the quarter the cost conversation opens the security conversation.

  4. Quarter 4

    Something surfaces, and the audit trail is the thing you needed six months ago

    Behavioral history cannot be reconstructed retrospectively. Whatever was not recorded at the time is simply gone.

The window to act is now, while the infrastructure is still forming

  • 33%

    of enterprise software applications will include agentic AI by 2028

    Gartner

  • 15%

    of day-to-day operations completed autonomously by 2028

    Gartner

  • 40%

    of agentic AI projects will be cancelled due to unmanaged risk

    Gartner

  • 12

    months in which every major platform shipped agent capability

    Anthropic, Microsoft, OpenAI, GitHub, AWS

Closing the agent governance gap starts with getting closer to the agent

Every requirement on this page points the same direction: understanding has to come from inside the agent’s own execution path, not from a boundary near it.