Home Projects Blog About Schedule a call DE

Shipping an LLM Feature to Production: What They Don't Tell You

The model is the easy part. Instrumentation, latency budgets, failure modes, and cost management are where production LLM features actually break — and where most teams are completely unprepared.

Latency is the UX

A 3-second API response is catastrophic on mobile. Users have been trained by instant-feedback interfaces — the moment there’s a perceptible pause, anxiety kicks in. For LLM features, where model latency is inherently higher, you need a latency strategy before you need a model strategy.

Streaming is the baseline. If your LLM API supports streaming responses, implement it first — don’t wait. A progressive display of the response transforms a 4-second wait into a 400ms time-to-first-token that feels immediate. Skeleton screens and typing indicators are supporting acts; streaming is the main event.

Set a latency budget. For our property valuation feature, the budget was 200ms for the inference path. We hit 97ms median. Anything that risked breaching 200ms needed a different architecture (edge deployment, smaller model, caching).

You Need LLM-Specific Instrumentation

Standard analytics events won’t capture what you need to know about an LLM feature. You need a separate event schema for AI interactions:

  • request_id — unique per LLM call for tracing
  • model_version — which model and which version
  • prompt_template_id — which prompt variant was used
  • input_tokens / output_tokens — for cost tracking
  • latency_ms — time from request to first token and to completion
  • user_feedback — thumbs up/down, edit, regenerate signals

The user feedback signal is gold. Track every regenerate click, every manual edit, every dismissal. That’s your ground truth for model quality — far more reliable than internal evaluations.

Define Your Failure Modes Before You Ship

LLMs fail in ways that traditional software doesn’t. They hallucinate, they return low-confidence answers with high confidence, and they can produce subtly wrong outputs that are harder to detect than crashes.

Pre-launch checklist

Before shipping any LLM feature: define what a bad output looks like, build a detection heuristic for it, and decide what the graceful fallback is. A good fallback is better than a confusing LLM response.

For the airport AI localization work, we defined failure not as an error response but as a response that fell outside the cultural register for the target market. That required human evaluation, not automated detection — plan for that overhead.

Cost is a First-Class Metric

LLM API costs scale with usage in a way that other infrastructure doesn’t. A feature that feels free in development can generate a £4,000/month bill in production. I’ve seen this blindside product teams repeatedly.

Instrument token usage from day one. Set up cost dashboards before you set up feature dashboards. Define a cost-per-session ceiling and alert when you approach it. Cache aggressively where input similarity is high — even approximate caching can reduce token spend by 30–40% for common query patterns.

The Prompt Is Production Code

Your prompt template determines output quality, cost, and failure rate. It deserves version control, testing, and a deployment process — not a quick edit in a config file. Use a prompt management system. Track which template version produced which outputs. Roll back when a prompt change degrades quality metrics.

A/B Testing LLM Features Is Different

Traditional A/B testing assumes deterministic output from a given input. LLMs are stochastic — the same input produces different outputs across calls. Your experiment design needs to account for this variance. Use large sample sizes, run tests for longer, and focus on outcome metrics (task completion, user edits, regeneration rate) rather than output similarity scores.

The good news: once you’ve built the instrumentation correctly, you have richer signal than any other feature type. User feedback, latency distribution, token spend, and downstream behaviour all combine into a picture of model health that no other analytics gives you.

Want this applied to your product?

Free 30-min data audit · No prep needed · Actionable gaps

Get your free audit →
Discussion

Leave a comment