Back to the blog

Running agents like operations

AgentsOperationsGovernance

Companies are past the point of having one agent. They have several, often in different departments, built by different people, on different platforms, with different ideas about what good looks like. Each one seemed manageable on its own. Together they are a fleet, and nobody is running it.

The shift we recommend is to stop thinking of agents as software features and start thinking of them as operations. Operations get managed: they have metrics, rhythms, owners, and a person who is accountable when they go wrong. Here is what that looks like applied to agents.

An inventory before anything else

You cannot run what you cannot see. The first job is a register: every agent in the organization, who owns it, what it does, which tools it can touch, what permissions it holds, and what it costs. Most organizations discover agents they did not know about during this step. Some discover agents with permissions nobody would have approved.

The register is a living document, and adding a new agent to it should be part of launching one. An agent not in the register is an agent not in the operation.

The metrics that matter

Operations run on a small number of numbers looked at regularly. For agents, four per agent and three for the fleet are enough.

Per agent: success rate against the owner's definition of good, override rate at the human gates, cost per run, and latency. These are the decay detectors. When one moves, someone looks.

For the fleet: total spend against budget, incident count and time to resolve, and the number of agents with a current owner and a passing evaluation. That last one is the health metric. An agent with no owner or a failing eval is a liability, however useful it was last year.

The rhythms

Weekly: each owner glances at their four numbers. Twenty minutes. Anything crossing a threshold becomes a ticket.

Monthly: the fleet review. Spend, incidents, decay signals, and the register reconciled against reality. Model updates from vendors are evaluated against each affected agent's harness before they are adopted, not after.

Quarterly: the portfolio question. Which agents are earning their keep, which should be extended, which should be retired. Agents accumulate. Some should not.

Governance as code, not as policy

The policies that govern agents, what they may do, what they may spend, what needs a human, are only real if the system enforces them. A written policy that an agent can exceed is a suggestion.

So the fleet needs a shared layer where those limits live: spend caps per agent, action allow-lists, approval gates for consequential actions, and an audit trail that captures every decision in a form a person can read. Individual agents inherit the policy rather than implementing their own version of it, which is how you get consistency across departments that never talk to each other.

The operator role

Someone has to run this. In a small company it is one person with a fixed slot in the week. In a larger one it is a function, often sitting with operations or IT rather than engineering, because the job is running things rather than building them.

The operator does not need to build agents. They need to hold the register, watch the fleet numbers, run the monthly review, handle incidents, and know when to escalate to whoever maintains the agents. That is the shape of the Agent Operations work we do for organizations with fleets: we act as that function, with the tooling, until the organization wants to hold it in-house.

The question this answers

The reason to do all of this is not tidiness. It is that leadership will eventually ask the question every operation gets asked: what are we running, what does it cost, is it working, and what happens when it fails? An organization that can answer in a page has a fleet. One that cannot has a pile of experiments with permissions.

Tell us where AI is stuck.

One conversation — we’ll tell you if we can help, and what we’d do first.

Book a call