RAG — retrieving your own documents and handing the relevant pieces to a model at answer time — remains the workhorse technique for making SI answer from your business rather than from the internet's average opinion. It is also the technique with the widest gap between a demo and a dependable system, because RAG fails politely: the model answers fluently either way, and only the customer knows the answer came from the wrong paragraph.
The uncomfortable rule of thumb: in a struggling RAG system, the model is almost never the problem. Retrieval is. If the right passage is in the context, current models will use it well; if the wrong passage is retrieved, no amount of prompting rescues the answer. So the reliability budget belongs upstream.
Your data is worse than you think
Every RAG project begins with the discovery that the business's knowledge is not a tidy corpus. It is a returns policy in three versions, a price list that contradicts the website, an FAQ written before the product changed, and tribal knowledge that was never written down at all. Indexing that pile faithfully gives you a system that faithfully retrieves contradictions.
- Nominate one canonical source per topic — one returns policy, one shipping table, one allergen list — and index only the canonical version.
- Delete or exclude superseded documents rather than hoping retrieval ranks the new one higher. Hope is not a ranking function.
- Write down the unwritten answers. The questions customers actually ask are in your inbox; the top twenty deserve a canonical paragraph each.
- Date-stamp everything, so freshness can be a retrieval signal rather than a mystery.
Chunk by meaning, not by size
Documents are indexed in chunks, and chunking is where retrieval quality is quietly decided. Split a refund policy mid-clause and the retrieved fragment says "within 30 days" without saying of what; the model then completes the thought on its own. The rule: a chunk should be a self-contained answer to some question — a whole policy clause, a whole product description, a whole FAQ pair — with enough header context attached that it still makes sense when it arrives alone.
Retrieval evals, run separately
Build a set of golden questions — real customer phrasing, not tidy test phrasing — each mapped to the passage that should be retrieved. Score retrieval on its own: was the right chunk in the top results? This isolates the layer that is actually failing.
Answer evals, run end-to-end
Then score final answers against what the business would say — correct, grounded, and honest about gaps. An answer eval without a retrieval eval tells you something is wrong; the pair tells you where.
Fifty golden questions maintained honestly beat five thousand generated ones. Run them on every change — new documents, new chunking, new embedding model — because RAG regressions are silent by nature, and the eval set is the only smoke alarm you get.
Freshness is a feature, staleness is an incident
A RAG system is a cache of your business, and caches go stale. The product sold out this morning; the delivery cutoff moved for the holiday; the price changed at noon. An agent quoting yesterday's truth with today's confidence is worse than one that says it does not know, because the customer has no way to tell the difference.
Two practices cover most of it. First, re-embed on change: when a document, product or policy is edited, its chunks are re-indexed then, not on a weekly schedule. Second, split facts from prose: volatile state — stock, prices, order status, availability — should never live in the document index at all. It belongs behind live lookups, fetched at answer time, so the document layer only ever holds the slow-moving truth.
Retrieval-augmented generation is a supply chain. The model is the last mile, and the last mile cannot fix what the warehouse shipped.
How this ships in a Phoxta business
Every Phoxta storefront runs this architecture per tenant. Each business's agent retrieves from that business's own knowledge — its policies, its product catalogue, its written answers — and never from a neighbouring tenant's. Volatile state comes from live reads of the actual order, booking and inventory records rather than from the index. Editing knowledge in the console re-indexes it, so the agent's answers move when the business moves.
And one honest limit to end on: RAG makes a model answer from your documents. It cannot make your documents agree with each other, and it cannot answer questions your business has never written down. The systems that feel intelligent are the ones sitting on knowledge someone curated — which is a job for an owner, not a model.




