Every agent demo looks the same: a model plans, calls a few tools, and finishes a task while the audience applauds. Production looks different, because production has customers, money and consequences. This is an inventory of what agentic SI genuinely does well in live businesses in 2026 — and the machinery that has to surround it before it is safe to leave running.
"Agentic" means one specific thing worth holding onto: the model does not just answer, it acts — it reads state, calls tools, and changes something in the world. That single step from talking to doing is where all the value and all the risk live, so it is where the engineering effort belongs.
Reading is solved; writing is where it gets serious
Give an agent read access to real state — orders, bookings, stock, a customer's own history — and it becomes reliably useful almost immediately. "Where is my order" answered with the actual tracking status, at any hour, on whatever channel the customer used, is the workhorse of agentic SI in commerce. It is unglamorous, high-volume, and it works because the agent is reporting facts it can see rather than generating plausible ones.
Read tools
Look up an order, check availability, quote a policy as written, summarise a conversation. Low blast radius: the worst outcome of a bad read is a wrong sentence, which the next message can correct.
Write tools
Issue a refund, move a booking, change an address, cancel an order. High blast radius: a wrong write moves money or breaks a promise, and no follow-up message un-moves it.
The systems that survive contact with production treat those two categories completely differently. Reads are given freely. Writes are governed: each kind of action is individually set to off, approve-first, or automatic, and the setting is a business decision made by the owner, not a default made by the vendor.
The approval gate is the technology
The most important component in a production agent is not the model. It is the queue between the agent's intention and the action's execution. When an overnight customer asks for a refund, the agent should be able to gather the order, check the policy, draft the resolution — and then park it in an approval queue with the full conversation attached, for a human to release with one tap in the morning. That pattern gets you most of the labour saving with almost none of the risk, and it is how trust is earned before any action graduates to automatic.
| Action | Sensible default | When to graduate to automatic |
|---|---|---|
| Answer from real order or booking state | Automatic | Day one — it is a read |
| Send a reply on the customer's channel | Automatic | Day one, with escalation rules |
| Move or amend a booking | Approve-first | After weeks of clean approvals |
| Issue a refund under a set amount | Approve-first | Small amounts, clear policy, audited |
| Change prices or delete records | Off | Rarely, if ever |
How production agents actually fail
The failure modes are by now well-catalogued, and none of them are exotic. An agent calls the right tool with the wrong argument — the correct customer, the wrong order. It answers from stale context because the underlying record changed mid-conversation. It loops, retrying a failing tool with growing confidence. Or it does the most human thing of all: states something false, fluently, because nothing in its context contradicted it.
- Wrong-argument tool calls — caught by validation on the tool side, never by trusting the model's formatting.
- Stale reads — caught by fetching state at answer time rather than conversation start.
- Loops and runaway sessions — caught by hard caps on steps, spend and time per conversation.
- Confident wrongness — caught by grounding answers in retrieved records and policies, and escalating when nothing grounds.
Notice that every one of those catches is boring infrastructure: validation, caps, logging, escalation. Production agentic SI is roughly one part model and three parts plumbing, and teams that resent the plumbing ship the incidents.
An agent you cannot audit is not an employee. It is a liability with an API key.
What this looks like when you buy it rather than build it
This is the architecture every Phoxta business ships with, because it is the only shape we trust in production. The customer-facing agent answers on web chat, SMS, WhatsApp and email with the real order and booking state behind it. The owner's Operator agent can act — but only through governed write-tools, with per-action policies, an approval queue, and an audit trail that records what was done and why. The exciting part of agentic SI is the acting. The trustworthy part is everything wrapped around it.
If you are evaluating any agent product, ask three questions: what can it write to, who approves, and where is the log. Vendors with good answers will show you screens. Vendors without them will show you demos.




