Skip to content

Field NotesWhat Is Managed Agent Operations?

AI Agents

What Is Managed Agent Operations?

Glyph-field title card on dark carbon: dense aiAgents texture glowing cyan, article title "What Is Managed Agent Operations? What You" on staggered dark slabs.
Managed agent operations is a service, not software. A provider runs your AI agents in live production while your business keeps the decisions. The provider owns monitoring, exception triage, workflow changes, integration upkeep, and incidents. You own policy, approval limits, and whether an agent gets to act at all. What you buy is operating capacity.

Essential Insights

  • "Managed agents" and "managed agent operations" are two different purchases. One is hosted plumbing for running agent code. The other is people who run the workload and answer for it on a Tuesday afternoon.
  • The things that kill agent projects are operational, not technical. Gartner's widely cited forecast names escalating costs, unclear business value, and inadequate risk controls, not weak models.
  • An AI agent is software that gets a goal, some access, and permission to take steps on its own. Access is the part that turns a demo into an operating risk.
  • Certain responsibilities can't be bought at any price. Approval limits, policy exceptions, and the definition of a good outcome stay with you no matter who runs the software.
  • Make a vendor fill in the responsibility split before you discuss price. Every blank row becomes your team's unpaid overtime.

Two different things are called "managed"

Search the phrase and you'll hit a product category first. In April 2026 Anthropic launched Managed Agents on Claude, which InfoQ described as a managed execution layer that separates agent logic from runtime concerns like orchestration, sandboxing, state management, and credentials. Useful thing. Also a hosting product. It runs the code; it has no opinion about whether your refund workflow should have fired.

Managed agent operations is the other thing: a service where a provider takes on the ongoing work of running agents against your real business, and stands behind the result. The hosting question is "where does this process execute." The operations question is "who gets paged when it executes wrong."

Founders conflate the two because vendors are happy to let them. A platform contract sounds like an outsourcing contract until the first bad week, when you discover the platform's obligation ended at uptime and everything past uptime was always yours.

So before comparing prices, work out which category the thing in front of you belongs to. A hosting product charges for compute and gives you a console. An operations service charges for coverage and gives you somebody to call. A software licence with a support inbox attached is still a software licence. The test is simple: ask what happens at 4pm on a Friday when the agent starts approving things it shouldn't. If the answer is a documentation link, you bought infrastructure.

What actually breaks after the demo

The demo is the easy part. Your agent drafts the reply, reconciles the invoice, updates the record, and everyone in the room nods. Then it meets production, which is where the invoice has a missing field, the customer exists twice in your CRM, the policy changed last week, and nobody updated the workflow.

Gartner's June 2025 prediction that more than 40% of agentic AI projects will be canceled by the end of 2027 got passed around as proof the whole category was hype. Read the reasons and it says something more specific: escalating costs, unclear business value, and inadequate risk controls. Model quality isn't on the list. Gartner also estimated that out of the thousands of vendors selling under the agentic label, only around 130 had genuine autonomous capability behind them, as Forbes reported in July 2026. The industry now calls the rest agent washing.

The gap between adoption and production is the number that should interest you. Forrester's 2026 assessment of the category, bluntly titled "Companies Are Chasing, Few Are Catching", found roughly three-quarters of enterprises adopting agentic AI and only a small fraction running it in real production. That's not a story about capability. Buying is easy. Operating is the bottleneck.

Two more findings sharpen the picture. The UK AI Safety Institute analyzed more than 177,000 agent tools built between late 2024 and early 2026, and found that action tools, the ones that send the email or move the money rather than describing it, climbed from 24% to 65% of usage across sixteen months. Meanwhile McKinsey's 2026 AI Trust Maturity Survey put average responsible-AI maturity at 2.3 out of 4, with only about 30% of organizations reaching level three or higher on governance and agent controls specifically. Agents are being handed the ability to act far faster than anyone is building the discipline to supervise it.

Managed agent operations exists because that discipline is a job. Somebody has to hold it. The only real question is whether that somebody is on your payroll.

Who does the work

Read the split before you read a price. Every row below is work that exists whether or not anyone is assigned to it.

Who owns each part of running an agent in production
The workWho owns itWhat it looks like when nobody owns it
Watching agents runThe providerYou learn a workflow broke from a customer complaint instead of a monitor
Routine exception triageThe providerA stuck queue grows quietly until someone notices the backlog
Prompt and workflow changesThe providerThe workflow drifts away from the policy it was built to enforce
Integration upkeepThe providerAn upstream field gets renamed and the agent fails without saying so
Incident containment and rollbackThe providerThe agent repeats the same mistake while people argue about who can stop it
Cost and usage controlThe providerToken spend climbs for months before anyone reads the invoice closely
Access scope for tools and dataThe provider proposes, you approveThe agent holds far more access than the job actually needs
Approval limits on consequential actionsYouMoney moves or customers get contacted with no human boundary in front of it
Policy exceptions and overridesYouEdge cases route to a provider who has to guess at your intent
The written definition of a good outcomeYouThe project has no agreed success measure and goes quiet at budget review
Whether the workload should be automated at allYouAn unused agent stays live, holding access and consuming budget

Marshal operates, agents work, clients approve. Any row a vendor leaves blank defaults to you, so get the blanks filled in before you get a price.

Below is the split. Read it as the actual job description of running agents in production, because every row is work that happens whether or not it's assigned.

The pattern worth internalizing: Marshal operates, agents work, clients approve. The provider does the running. The agent does the task. You hold the judgment. Any arrangement that blurs those three ends up with either an agent making business decisions it has no standing to make, or a founder doing systems administration at 11pm.

Notice which rows sit on your side. They're short, and they're all judgment: what the limits are, what the exceptions mean, what counts as success. That's the correct division of labour, and it's also the honest one. Nobody can sell you a service that decides how much financial risk your business will accept from an automated action. What they can sell you is everything else, which happens to be the bulk of the hours.

The ongoing work is genuinely ongoing

Founders underestimate the operations bill because launch feels like the finish line. Microsoft's cloud adoption guidance for managing agents across an organization reads like a maintenance manual, and it recommends things like regular audits of every deployed agent, retiring the ones nobody uses, running a single control plane rather than administering agents one at a time, and putting caps on token spend and request rates so costs can't run away. Its warning about ungoverned rollout is specific: shadow AI spreading through the company, unpredictable budget overruns, and dormant agents quietly expanding your attack surface.

None of this is exotic engineering. It's housekeeping, and housekeeping is exactly what gets skipped when the person responsible also runs sales. Three items cost the most in practice:

  • Integration upkeep. Your agent reaches into a CRM, an inbox, an accounting system, a scheduler. Those systems change under you. A renamed field or a tightened permission breaks a workflow, sometimes loudly, often silently.
  • Exception triage. Real work throws edge cases weekly. Somebody has to look at what stalled, decide whether it stalled correctly, and either clear it or change the rule.
  • Change management. Every policy change in your business is a change to the agent's instructions. Miss the connection and the agent keeps enforcing last quarter's policy with perfect consistency.

The judgment you can't hand over

Here's the part vendors soften: buying managed operations doesn't zero out your involvement, and you should be suspicious of anyone who says it does. Four things stay yours.

Approval limits. You decide the point at which an action needs a human. Refunds under $200 go through, over $200 come to you. That threshold is a business risk decision, and no provider can set it for you.

Policy exceptions. When a case falls outside the rules, the answer comes from your intent, not from inference. A provider guessing at your intent is a provider generating your next incident.

What counts as good. The written success measure, and the person who agreed to it. Forbes framed it as the question an executive should be able to answer before greenlighting anything: what's the success metric, who agreed to it, and how fast can someone roll the thing back. Projects without that answer are the ones that go quiet at budget review.

Whether the workload should exist. Some jobs shouldn't be handed to an agent yet. Saying no is your call, and it's frequently the right one.

Calibrated control is the goal, not maximum control. An agent that needs sign-off on every trivial step hasn't saved anyone anything; it's a manual process wearing automation as a costume. Put the checkpoints where a mistake is expensive and leave the cheap steps alone.

What to demand before you sign

Ask for these in writing. A serious provider has answers ready, and the shape of the hesitation tells you plenty.

  • The named workloads in scope, described as jobs your business recognizes, not as features.
  • The approval boundary: which actions execute freely, which wait for a person, and who that person is.
  • Monitoring you can actually see, not a quarterly summary of how well things went.
  • The change process when your policy shifts, including how fast a change lands.
  • A stop switch, named: who can halt an agent mid-task, and how long it takes.
  • Full action logs available on demand, so you can reconstruct why an agent did something six months later without an archaeology project.
  • Exit terms. What you keep if you leave: workflows, logs, configuration, access records.

That last one matters more than founders expect. Regulatory pressure is heading toward documented human oversight of higher-risk systems, and reporting on the European Union's 2026 Digital Omnibus agreement indicates the AI Act's related compliance deadline moved out to December 2027. Whatever the eventual enforcement, records that live only inside a vendor's console are records you don't control.

The honest tradeoff

Managed agent operations transfers labor and expertise, not accountability. You still own the risk, because you own the business. What you're buying is somebody competent standing between your operations and the failure modes that killed 40% of the projects in Gartner's forecast, plus the time you get back by not becoming an amateur operator of software you didn't build.

Priced correctly, it beats hiring for it. A capable internal owner for this work is a real salary, and one person doesn't cover holidays or leave. Priced badly, or scoped vaguely, it's a monthly invoice for a dashboard you never open. The responsibility split is what separates the two, which is why you should get it filled in before anyone quotes you a number.

One more test, cheap and revealing: ask a prospective provider to walk you through a failure they handled on somebody else's workload. Not a success story, a failure. What broke, how they found out, what they changed, how long the customer sat exposed. Anyone who has genuinely operated agents in production has three of those stories and tells them without flinching. Anyone who has only sold agents will redirect you to a case study. The category is young enough that operating experience is scarce and easy to fake, so the diligence that matters is less about the technology than about whether the people quoting you have ever been on the wrong end of it.

Frequently Asked Questions

Is managed agent operations the same as an AI consultancy?

Managed agent operations is different from consulting because consulting ends with a recommendation and a bill. Managed operations means the provider is still there in month seven, watching the workload run, fixing the integration that broke, and handling exceptions. Consulting sells thinking. This sells running.

How much of my time will this still take?

Expect to spend real time on approvals and exception decisions, especially in the first weeks while the boundaries get set. That time should shrink as the rules harden, but it never hits zero, because approval limits and policy calls are yours by definition. Any provider promising zero involvement is describing a product they haven't operated.

What if the agent does something wrong?

Ask that question during procurement, and get the answer in writing. A workable arrangement names who detects the problem, who can stop the agent mid-task, how the action gets reversed where reversal is possible, and what the provider owes you. Gartner's cancellation reasons include inadequate risk controls, and this is the control most often absent.

Can a small business justify this?

Small businesses often have the clearest case, because they lack an operations team to absorb the work. A ten-person company running one billing agent still needs monitoring, upkeep, and someone on call. What decides it isn't headcount but whether the workload has enough repeated volume to be worth automating and enough consequence to be worth supervising.

How do I know it's working?

Pick the measure before launch and keep it boring: volume handled without a human touch, exception rate, time from exception to resolution, and cost per completed job. Compare against how the work ran before. If a provider can't report those numbers on demand, they aren't operating the workload, they're hosting it.

Reimagine your work-lifewith managed agents on the job.

Business, at the speed of AI.