Agent Operations: The Role Without a Title Yet

Someone in every agent-running organization does this job without being hired for it. It has no name, no allocation, and no incident that forces it — until whoever was doing it informally leaves.

Someone in every organization running agents seriously ends up doing a job nobody hired for. They watch whether the agents are still working, chase down why quality drifted, keep the eval set current, notice the cost went up, and handle it when something goes wrong.

It's a real role. It has no name, no job description, and usually no allocation — which means it gets done in the gaps by whoever cares most, until they stop.

What the work actually is

Watching for silent degradation. Agents don't fail loudly. Turns per task creeping up, escalation rate shifting, a tool's error rate rising, cache hit rate dropping. Someone has to look, because nothing pages.

Maintaining the eval set. Traffic shifts, and a suite built on last quarter's requests stops representing production. Adding cases from real failures, retiring dead ones, re-sampling periodically.

Running model migrations. A provider deprecates a version, or a new one arrives. Comparing on a frozen eval set, reading diverging cases, testing whether workarounds are still needed, recalibrating thresholds.

Cost management. Watching cost per completed task, finding the runs that dominate spend, deciding where routing or context trimming pays.

Incident response. When an agent does something wrong at volume, someone reconstructs what happened, determines the blast radius, and decides what to roll back.

Sampling for quality. The weekly read of actual outputs, which is the only mechanism that catches quiet wrongness and the one that doesn't scale.

Prompt and tool maintenance. Pruning accumulated instructions, fixing tool descriptions when selection drifts, keeping the surface coherent.

Why it doesn't get staffed

No incident forces it. The work prevents slow degradation, and slow degradation produces no event that triggers a hiring decision.

It looks like several existing jobs. Parts resemble SRE, parts data science, parts QA, parts product. Each team assumes another owns it.

Its output is absence. A well-operated agent behaves normally, which looks like nothing happening.

⚠️ The predictable result: it's done informally by whoever built the thing, until they move on, and then it isn't done at all. The agent degrades quietly for months.

✅ What staffing it looks like

A named owner with a number. Completion rate, cost per completed task, or escalation rate — someone accountable for it and able to affect it. Unowned quality drifts.

A scheduled review, not an intended one. Weekly: read the dashboard, sample a dozen outputs, check the escalation queue. An hour, on the calendar.

A defined migration process. So a model deprecation is a routine project rather than a scramble.

On-call coverage with real authority — able to roll back a prompt, disable a tool, flip a kill switch without a deploy.

Budget for the eval set. Its maintenance is the work that keeps everything else measurable, and it's the first thing dropped under pressure.

💡 Why it's a good role to take

For anyone positioning deliberately, this is an unusually good opportunity:

  • It's unclaimed, so it's available without a title change.
  • It requires the durable skills — verification, systems thinking, domain knowledge, judgment about what matters.
  • It produces visible evidence of judgment: the regression you caught, the migration that went smoothly, the cost you halved.
  • It becomes a named role eventually, and the person already doing it is the one who gets it.

The pattern is familiar — SRE existed as informal work before it had a name, and the people doing it informally defined it.

🔍 Whether your organization needs it yet

  • Can anyone say what your agents' completion rate was last month, and whether it moved?
  • When did someone last read a sample of real outputs?
  • Who would handle a model deprecation notice?
  • Is the eval set current with this quarter's traffic?
  • Do you know cost per completed task, per feature?

Two or more "nobody knows" answers means the work is already needed and already not being done.

The takeaway

Agent operations is watching for silent degradation, maintaining the eval set, running migrations, managing cost, responding to incidents, and sampling outputs. It's real work that no incident forces and that looks like four other jobs, so it goes unstaffed until whoever was doing it informally leaves. Name an owner with a number, put the review on the calendar, and give on-call real authority — and if you're looking for a durable position, this one is unclaimed and made of exactly the skills that are appreciating.

Keep reading

Similar posts

Matched on shared tags and category — the more bars, the stronger the overlap with what you just read.