Choosing the right LLM for your use case

Phoxta
July 17, 2026 · 8 min read
SHARE THIS ARTICLE
Choosing the right LLM for your use case
The question "which LLM should we use" contains a wrong assumption: that a business uses one. In practice a production system is a set of jobs — routing, extraction, drafting, reasoning — and the honest answer is a routing policy: which tier of model handles which job, and what evidence would change the assignment. Framed that way, most of the decision makes itself.

The market in 2026 has settled into recognisable tiers, whatever the logos on them. What follows deliberately avoids naming models — the leaderboard changes quarterly, but the shape of the trade-off has been stable for years and is the part worth learning.

The tiers, and what they are actually for
TierBest atLatencyRelative costTypical jobs
Small / fastClassification, routing, extractionSub-second1×Intent detection, tagging, form-filling
Mid-tierGrounded drafting and conversation1–3 s5–15×Customer replies, summaries, RAG answers
FrontierMulti-step reasoning, tool orchestration3–10 s+30–100×Agent planning, hard escalations, analysis
Capability tiers as they trade off in practice. Costs are relative, not quoted — the ratios are the stable part.

The instinct to reach for the frontier tier for everything is the most expensive habit in applied SI. Frontier models are remarkable at the hardest step of a workflow — and indistinguishable from mid-tier models on the easy steps, except on the invoice and the stopwatch.

When small models win

A surprising share of real workload is not generation at all. It is deciding: which mailbox does this message belong in, is this a refund request, what is the order number in this sentence, is this conversation angry. These are classification and extraction jobs, they run thousands of times a day, and a small model does them in a few hundred milliseconds at a fraction of the cost — often more consistently than a large model, which is tempted to editorialise.

  • High-volume, low-ambiguity, checkable output — the small-model sweet spot.
  • Anything on the critical latency path, where the customer is watching a typing indicator.
  • Structured output jobs, where the schema does the thinking and the model does the filling.
  • First-pass triage, with the genuinely hard cases routed up a tier — the cascade pattern.
Latency is a product feature

Chat feels broken beyond a few seconds; voice is unusable beyond one. Latency budgets, not benchmark scores, disqualify most models for most interactive jobs — which is a reason to route, not a reason to compromise on the hard steps.

Context is not memory

Million-token windows tempt you to shovel everything in. Retrieval that selects the right two thousand tokens beats a full window on cost, speed and, frequently, accuracy — attention is not evenly spread across a haystack.

Switching is cheap if you built for it

Model choice is a decision you will make repeatedly, because the market moves under you. Two practices keep it cheap. Keep the plumbing model-agnostic, so a model is a configuration entry rather than an architecture. And keep an eval set — your own questions, your own tone, your own edge cases — so that when a new model ships, the evaluation is an afternoon against your yardstick rather than a fortnight against a press release.

Benchmarks tell you what a model can do. Only your own evals tell you what it will do with your customers, your policies and your worst Tuesday.

How Phoxta routes, and why you mostly should not care

Phoxta businesses run exactly this portfolio pattern under the hood: fast models on classification and routing, stronger models on customer-facing answers and on the Operator agent's multi-step work, with retrieval keeping context small and token usage metered per business so cost stays a visible number. Owners never pick a model, and that is the point — model selection is an operations decision that should be made once, measured continuously, and revised quietly.

If you are making the choice for your own stack, resist ending the meeting with a model name. End it with a routing table, a latency budget and an eval set. The name in each row will change; the rows will not.

Your privacy choices

We use essential storage to keep you signed in, secure your account and remember your choices. Optional analytics helps us understand how Phoxta is used.