There is a moment in every SI deployment when the demo is over, the agent is live, and someone asks a very ordinary question: what did it do yesterday? If the honest answer is "we are not entirely sure", you do not have an operations problem — you have an operations absence. AgentOps is the unglamorous discipline of being sure.
The good news for business operators: this is not research. The practices are settled, they are mostly borrowed from twenty years of running ordinary software, and the right time to adopt them is before the first incident rather than after it.
Tracing: every conversation is a flight recorder
The unit of observability for an agent is the trace: the full record of one conversation — what the customer said, what the agent retrieved, which tools it called with which arguments, what came back, and what it finally did or said. When a customer disputes an outcome, the trace is the difference between an answer and an argument. When behaviour drifts after a model or prompt change, traces are how you see it before your customers describe it to you.
- Log every tool call with its arguments and result — the actions matter more than the words.
- Keep traces reviewable by a human in one screen; a trace nobody reads is storage, not observability.
- Sample and review a handful of ordinary conversations weekly, not just the escalations — drift starts in the ordinary ones.
- Record why an action was allowed: which policy setting, whose approval. Auditors and refund disputes both ask this question.
Guardrails live at three layers
Input guardrails decide what reaches the agent: rate limits, spam filtering, and treating retrieved or customer-supplied text as data rather than instructions — the standing defence against prompt injection. Output guardrails decide what leaves: no invented discounts, no promises the policy does not make, escalate when ungrounded. Action guardrails are the layer that matters most and is skipped most: per-action write policies, amount thresholds, and an approval queue between intention and execution.
Cost budgets
Token spend is a real line item and an early-warning signal at once. Meter usage per conversation and per tenant, cap the runaway cases, and alert on the trend — a cost spike is usually a loop, a prompt regression or an attack, discovered by the finance graph.
Latency budgets
A brilliant answer in ninety seconds is an abandoned chat. Set a budget per reply, measure the slowest tenth rather than the average, and spend the budget deliberately — retrieval, tool calls and drafting all bill against the same clock.
Red-team your own agent before someone else does
Every public-facing agent gets probed — for free merchandise, for other customers' data, for a jailbreak that makes a screenshot. Run the attacks yourself first, calmly and on a schedule: ask it to reveal another customer's order, instruct it to ignore its rules inside a pasted "policy", talk it toward a refund it should not grant. Each attempt that fails is a regression test; each one that succeeds is a fix you got to make privately.
| Signal | Watch for | What it usually means |
|---|---|---|
| First-response time, whole clock | Creeping upward | Tool or retrieval latency, not the model |
| Resolution without human handoff | Sudden drops | A knowledge gap or a broken tool |
| Escalations and approvals per day | Spikes | New question pattern, or a policy set too tight |
| Cost per conversation | Outliers and trend | Loops, prompt bloat, or abuse |
| Actions taken, by type | Anything unexpected | The report you read before it becomes an incident |
You would not employ a person whose work you never reviewed, whose spending you never saw, and whose decisions left no record. An agent deserves the same management, minus the awkward conversations.
What this looks like on a Phoxta business
We ship this discipline as standard equipment because operators should not have to assemble it. Every agent conversation on a Phoxta storefront is a reviewable record; every write-action goes through the governed policy layer — off, approve-first, or automatic, per action — with the approval queue and audit trail attached; token usage is metered per business, so the cost of the agent is a number on a page rather than a surprise on an invoice.
The pattern to hold onto: autonomy is not a switch, it is a dial, and observability is what earns each turn of it. Agents graduate from approve-first to automatic on evidence — weeks of clean traces, not weeks of good luck.




