Skip to main content
Blog

The Golden Catalog Is a Comforting Fiction

Pre-approved agent template libraries optimize for control over a static asset, when the thing they're trying to control isn't static at all.

Why pre-approved template libraries can’t keep pace with how fast agentic AI actually changes

Across medium and large enterprises, a particular pattern keeps showing up in how organizations are trying to get their arms around agentic AI. Call it the golden catalog: a curated, internally built set of templated agents, pre-approved and pre-blessed, that the business can roll out without leaning on external agent builders or third-party harnesses. The logic is understandable. If you build the catalog yourself, you control the surface area. You know what’s in it. You can point to internal sign-off and call the problem solved.

The trouble is that this approach solves for the wrong variable. It optimizes for control over a static asset when the thing it’s trying to control isn’t static at all.

Why the catalog falls behind the harness

The first problem is pace. The catalog model assumes agent capability is something you can freeze, test, and ship on a devsecops cadence: build, harden, release, repeat. That’s a reasonable operating model for software that changes at the speed software used to change at, but not for a category where the underlying harnesses (Claude, ChatGPT and Codex, Cursor) are shipping step changes in capability on a cycle measured in weeks, not quarters. Every major platform has shipped a step change in agent capability within the last twelve months. By the time a template has cleared the internal review and testing gates an enterprise rightly insists on, the model or harness it was built against has often moved on. The catalog ends up permanently behind, and the gap between what’s approved and what’s actually capable keeps widening.

The behavior gap A concept diagram plotting capability in production over time. Harness capability climbs in frequent steps and keeps rising; catalog approval status climbs in fewer, smaller steps and falls further behind. The widening space between the two lines is labelled "the gap". The behavior gap Harness capability vs. catalog approval status, over time time → capability in production harness capability catalog approval status the gap Illustrative - a conceptual sketch of the pattern, not measured data.
Fig 1: The behavior gap - a conceptual sketch of harness capability pulling away from catalog approval status over time

What’s needed is an operating posture that assumes continuous drift as the baseline condition, not an occasional exception to manage around.

The generic-versus-specific trap

The second problem is harder to fix with process alone, because it’s a design tension rather than a scheduling one. A template has to be generic enough to cover the range of business use cases a catalog is meant to serve, but specific enough that a non-technical end user can actually get value from it without needing to understand what’s happening underneath. Those two goals pull against each other. Too generic, and the agent becomes a shell that requires the kind of prompt literacy most business users don’t have and shouldn’t need. Too specific, and you’re back to building and maintaining a template per use case, which is the exact iteration burden the catalog was meant to avoid. In practice, a lot of teams resolve this tension by quietly retreating: the “agent” becomes a chatbot with a system prompt, and the organization ends up right back where it started, just with more infrastructure around it.

Templates are artifacts, behavior is the thing to govern

What both of these problems point to is that the golden catalog is trying to answer an infrastructure question using a product-development answer. Templates are artifacts. What enterprises actually need is a way to track and shape agent behavior as agents operate, adapt, and get swapped out underneath the templates that wrap them: visibility into what an agent is actually doing in a workflow, how its behavior drifts as the underlying model changes, and where its actions sit relative to the risk posture the business has defined. That’s a posture and observability problem that no amount of internal build discipline can turn into a template-versioning one.

Two models, one agent estate A five-row comparison table. The golden catalog applies to templates versioned like software, assumes capability can be frozen and shipped, runs on a quarterly build-test-ship cadence, breaks first at the harness it was built against, and produces a snapshot that ages the moment it ships. Behavioral assurance applies to agent behavior inside the workflow, assumes capability is continuously in motion, runs on continuous observability, breaks at nothing because it moves with the drift instead, and produces assurance over how agents behave now. Two models, one agent estate Where the two approaches actually differ The golden catalog Behavioral assurance What it applies to Templates, versioned likesoftware Agent behavior inside theworkflow What it assumes Capability can be frozen andshipped Capability is continuously inmotion Cadence Build, test, ship on aquarterly cycle Continuous observability What breaks first The harness it was builtagainst Nothing. It moves with thedrift instead. What it produces A snapshot that ages themoment it ships Assurance over how agentsbehave now Two operating models for the same agent estate - not a checklist to complete once.
Figure 2: Two models, one agent estate - a comparison of the golden catalog against behavioral assurance

Owning your agent estate is not owning a template library

None of this is an argument that self-building is misguided as an instinct. Wanting to own your agent estate rather than hand it entirely to a vendor is a sound instinct, and external harnesses and agent-builders aren’t going away - nor should they. But ownership of an agent estate shouldn’t be confused with ownership of a fixed set of approved templates. The former is about maintaining assurance over how agents behave as operational actors embedded in real workflows. It requires infrastructure that scales with how fast the category is actually moving. The latter is a snapshot of a moving target, and snapshots age the moment they’re taken.

The organizations that get this right will be the ones that treated agentic capability as inherently in motion from the start, building an operating model that moves with it rather than trying to hold it still.

Knowing what's in the catalog is the easy half. See how Agent Discovery finds the agents nobody registered and maps what each one is wired to, or talk to our team about the estate you already have.

Keep reading

Deep Research Discovery & Posture

A Guide to Agent Control

Geordie’s Field Guide shows the control points Anthropic, OpenAI, AWS, and Microsoft have built in natively – and where each falls short.