Phase 5: Production & Deployment

Latency, token usage, error rates & user satisfaction metrics

Intermediate ~14 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

When you build amazing smart computer programs, like the ones that can chat with you or help you learn, you want to make sure they're always working great for everyone. It's a bit like baking a bunch of cakes for a big party. You don't just put them in the oven and hope for the best; you need to keep an eye on things to make sure every single cake is perfect and makes people happy.

First, think about how long it takes for your cakes to bake. Some might be ready in 30 minutes, but what if one takes an hour because the oven got a bit slow? In the world of computers, we call this "latency" – how fast your smart program gives an answer. If it's too slow, people get bored waiting. You also need to watch your ingredients. Just like flour and sugar cost money, the "tokens" (tiny pieces of information) your program uses to think cost money too, and there's only so much memory it can use at once. You want to use just enough ingredients for a delicious cake, not too much or too little.

Next, you need to know if the cakes actually taste good! Sometimes a cake comes out burnt, or it's raw in the middle – those are clear problems. But what if it looks fine but just tastes really bland? Your smart program can make mistakes too, from giving completely wrong answers to answers that are just a little bit off. We call these "errors." Sometimes, people might even stop telling you the cake is bad because they've already decided not to eat from your bakery again – that's a silent problem you still need to fix!

And finally, are people happy when they eat your cake? Did they smile and ask for another slice, or did they just pick at it and leave it on the plate? This is how satisfied people are with your smart program. Maybe they tell you directly, "This is great!", or maybe you just notice they keep coming back to use it. When you build smart programs, your job is to watch all these things together – how fast it works, how many resources it uses, if it makes mistakes, and if people enjoy it – so you can make sure every "cake" your program delivers is a hit!

The four metric categories in this lesson map to four different failure modes. Latency failures hurt perception and SLA compliance. Token usage failures hurt your wallet and context window constraints. Error rate failures hurt reliability. Satisfaction failures hurt retention. Each requires a different instrumentation strategy and a different alert threshold.

Latency: measure at the right granularity

Average latency is nearly useless for AI systems. A single slow request in a batch can double your mean without touching the median. What you want is p50 (the typical request), p90 (the experience for the top 10% most affected users), and p99 (the worst 1%). For interactive chat, p99 above 5-6 seconds is a hard UX problem. For background summarization pipelines, p99 of 30 seconds may be fine. The threshold depends on your product contract, not a universal standard.

Break latency down further: time-to-first-token (TTFT) and tokens-per-second (TPS). TTFT measures how long before the user sees anything, which dominates perceived responsiveness in streaming UIs. TPS measures throughput once generation starts. A high TTFT with high TPS suggests queuing or cold-start issues at the provider. A low TTFT with low TPS suggests the model is generating slowly, which might mean you're on an overloaded tier. Tag latency by model, prompt template version, and endpoint so you can isolate regressions when you ship changes.

Token usage: cost and context window are two separate concerns

You already know tokens drive cost from the tokenization lesson. The production angle is that token usage needs to be sliced by feature, user cohort, and prompt version, not just totaled per day. If your /summarize endpoint averages 1,200 input tokens per call and your /chat endpoint averages 300, a 20% traffic shift toward summarization changes your cost structure significantly. Without per-feature tracking, you'll see the cost spike and have no idea where it came from.

The second concern is context window utilization. If your average request is consuming 85% of the context window, you're one slightly longer document away from a truncation error. Track the ratio of tokens used to the model's limit per call. Alerting when p90 utilization exceeds 80% gives you runway to fix chunking or routing logic before users hit errors.

Error rates: hard vs. soft failures

Hard errors are easy: HTTP 4xx/5xx, timeouts, JSON parse failures when you expected structured output, and finish_reason == "length" (the model ran out of tokens mid-response). These are straightforward to count and alert on. A hard error rate above 0.5% in a production LLM integration is worth investigating immediately.

Soft errors are harder and more important. A soft error is when the model returns a valid HTTP 200 with plausible-looking text, but the response is wrong. This includes hallucinated facts, format violations that slip past your parser, refusals that weren't triggered by a safety violation but by a vague prompt, and responses that are technically correct but miss the user's intent. Detecting soft errors requires either an LLM-as-judge layer (covered in aidev-evaluation-llm-judge) or heuristic classifiers: does the response contain a JSON object? Does it mention the entity the user asked about? Is the response longer than five words? These weak checks catch the worst cases cheaply.

Also track finish_reason distribution. In a healthy system, the vast majority of completions should finish with "stop". A rising share of "length" finishes means your prompts or user inputs are growing, and you're silently truncating responses. Users experiencing truncated responses rarely complain, they just leave.

User satisfaction: explicit and implicit

Thumb ratings and CSAT surveys have high value but low volume. Most users won't rate a response unless it was notably good or notably bad. That selection bias skews your scores. Treat explicit ratings as a high-signal sample, not a representative measure.

Implicit signals fill in the gaps. Retry rate (did the user immediately resend a similar query?) is a strong negative signal. Copy-paste detection (did the user copy the response to clipboard?) is a positive engagement signal. Follow-up clarification rate (did the user ask a follow-up question to fix the response?) often indicates the first response was unsatisfactory. Session abandonment after an AI response is a soft negative signal. None of these are perfect, but together they give you a satisfaction proxy that correlates well with explicit ratings and has 100x the volume.

Tradeoffs and scale

At 10 users, log everything to a file and query it manually. At 10,000 users, you need a structured logging pipeline (covered in aidev-monitoring-logging) feeding a dashboard. At 10 million users, sampling becomes necessary. You cannot afford to run an LLM judge on every response at scale, so you run it on a 1-5% sample, stratified by endpoint and user segment. The key is that your sampling strategy should be deterministic and reproducible so you can rerun analysis on the same slice. Latency and token usage can be logged 100% of the time cheaply because they're just numbers. LLM-based quality evaluation is expensive and should be sampled.

Key Takeaways

  • Track p50, p90, and p99 latency separately — averages hide the tail behavior that kills UX.
  • Tie token usage to individual features and users, not just aggregate totals.
  • Distinguish hard errors (HTTP 500, timeout) from soft errors (bad format, hallucination, refusal).
  • Implicit satisfaction signals like retry rate and copy-paste behavior often outperform thumbs up/down.

Pro tips

  • Track context window utilization (tokens_used / model_limit) per call, not just token count. When your p90 hits 80% of the window, you have a chunking or routing problem waiting to surface as truncation errors.
  • Log finish_reason on every call. A creeping increase in 'length' finishes is a silent regression — responses are being cut off, users are getting incomplete answers, and your error rate counter shows zero because HTTP 200 came back.
  • Instrument TTFT (time-to-first-token) separately from total latency, especially in streaming apps. Users tolerate long total generation times much better when they see tokens start appearing quickly. If TTFT is consistently above 1.5s, investigate provider queuing or cold-start issues before blaming generation speed.
  • When you A/B test prompt changes, segment your satisfaction signals by variant from day one. Post-hoc analysis that tries to correlate a deployment timestamp with a shift in thumb ratings is almost always confounded by traffic pattern changes or other concurrent experiments.

Common pitfalls

  • Mistake: Alerting on average latency instead of p99. Fix: Always set SLA alerts on p95 or p99; average latency masks tail behavior that causes the most user-visible failures.
  • Mistake: Tracking only total token usage per day. Fix: Tag every call with feature name and user segment so you can pinpoint which endpoint or cohort is driving cost spikes.
  • Mistake: Treating all errors as hard errors and ignoring soft failures like truncated responses or format violations. Fix: Check finish_reason and run lightweight heuristics on output structure to count soft errors separately.
  • Mistake: Relying solely on explicit thumb ratings for satisfaction data. Fix: Instrument implicit signals like retry rate and session abandonment alongside explicit ratings to get statistically meaningful volume.

Which satisfaction signal to prioritize

Option Use when Avoid when
Explicit ratings (thumbs, CSAT) You have a UI surface where ratings fit naturally and you need high-signal labeled data for model evaluation. Your user base is small or the interaction is high-frequency; response rates will be too low to be statistically useful.
Implicit behavioral signals (retry rate, copy-paste, abandonment) You need high-volume satisfaction proxies with no UI changes; works in any product surface. You need to distinguish between 'user rephrased for clarity' and 'user was dissatisfied'; behavioral signals are ambiguous.
LLM-as-judge quality scoring You need automated quality scores on sampled traffic to catch regressions after prompt or model changes. Evaluating every request in real time; cost and latency make it impractical above small traffic volumes without sampling.
Task completion rate Your AI feature has a defined goal (form filled, booking confirmed, search query resolved) you can instrument. The task boundary is fuzzy (open-ended chat, creative writing) where 'completion' is undefined.

Code Example

python
# openai>=1.0.0, requires OPENAI_API_KEY in environment
import time, os
from openai import OpenAI

client = OpenAI()

def call_with_metrics(prompt: str) -> dict:
    start = time.perf_counter()
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
    )
    latency_ms = (time.perf_counter() - start) * 1000

    usage = response.usage
    return {
        "latency_ms": round(latency_ms, 1),
        "prompt_tokens": usage.prompt_tokens,
        "completion_tokens": usage.completion_tokens,
        "total_tokens": usage.total_tokens,
        "content": response.choices[0].message.content,
        "finish_reason": response.choices[0].finish_reason,
    }

result = call_with_metrics("Summarize the water cycle in one sentence.")
print(result)

How this code works

This code’s main job is to send a request to an AI model and meticulously record important performance and cost metrics. It starts by setting up the OpenAI client, which requires an OPENAI_API_KEY to connect to the AI service. The core logic resides within the call_with_metrics function. This function uses time.perf_counter() to accurately mark the beginning of an AI interaction, ensuring precise measurement of how long the AI takes to respond before client.chat.completions.create is invoked.

Inside call_with_metrics, the actual AI call happens using client.chat.completions.create. This specifies the model to use (like gpt-4o-mini, a common choice for its balance of cost and speed) and the messages containing the prompt. Immediately after the response, time.perf_counter() is called again to calculate latency_ms, representing the time taken in milliseconds. The code then extracts detailed token usage information from response.usage, including prompt_tokens (input), completion_tokens (output), and total_tokens. Finally, it returns a neat dictionary containing all these metrics along with the AI’s generated content and the finish_reason for the response, providing a comprehensive overview of the AI interaction.

Production-grade example

Adds retries with backoff, per-feature structured logs, timeout handling, and soft-error detection.

python
# openai>=1.0.0, structlog>=23.0.0
import time, os, asyncio
from openai import AsyncOpenAI, APITimeoutError, RateLimitError, APIStatusError
import structlog

log = structlog.get_logger()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)

MAX_RETRIES = 3
BASE_BACKOFF = 1.0  # seconds

async def call_with_full_metrics(
    prompt: str,
    model: str = "gpt-4o-mini",
    feature: str = "unknown",
    user_id: str = "anonymous",
) -> dict:
    attempt = 0
    while attempt < MAX_RETRIES:
        start = time.perf_counter()
        try:
            response = await client.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": prompt}],
            )
            latency_ms = round((time.perf_counter() - start) * 1000, 1)
            usage = response.usage
            finish_reason = response.choices[0].finish_reason

            # Detect soft error: truncated response
            soft_error = finish_reason == "length"

            log.info(
                "llm_call_success",
                model=model,
                feature=feature,
                user_id=user_id,
                latency_ms=latency_ms,
                prompt_tokens=usage.prompt_tokens,
                completion_tokens=usage.completion_tokens,
                total_tokens=usage.total_tokens,
                finish_reason=finish_reason,
                soft_error=soft_error,
                attempt=attempt + 1,
            )
            return {
                "content": response.choices[0].message.content,
                "latency_ms": latency_ms,
                "total_tokens": usage.total_tokens,
                "soft_error": soft_error,
            }

        except RateLimitError:
            wait = BASE_BACKOFF * (2 ** attempt)
            log.warning("llm_rate_limited", feature=feature, attempt=attempt + 1, retry_in=wait)
            await asyncio.sleep(wait)
            attempt += 1

        except APITimeoutError:
            latency_ms = round((time.perf_counter() - start) * 1000, 1)
            log.error("llm_timeout", feature=feature, latency_ms=latency_ms, attempt=attempt + 1)
            attempt += 1

        except APIStatusError as e:
            log.error("llm_api_error", feature=feature, status_code=e.status_code, message=str(e))
            raise  # non-retryable

    log.error("llm_max_retries_exceeded", feature=feature, max_retries=MAX_RETRIES)
    raise RuntimeError(f"LLM call failed after {MAX_RETRIES} attempts")

How this code works

This code defines a robust way to interact with OpenAI's language models, specifically designed to gather critical performance data for AI observability lessons. The primary function, call_with_full_metrics, wraps an OpenAI API call, focusing on tracking key metrics like latency_ms, prompt_tokens, completion_tokens, and the finish_reason to understand model behavior. It ensures reliability by implementing a retry mechanism with exponential backoff for RateLimitError exceptions, and it logs APITimeoutError events, reflecting common real-world challenges. Structured logging with structlog records all these details, providing a rich dataset for analysis in an AI monitoring system.

Beyond simple success/failure, the code subtly captures a soft_error when the finish_reason is length, indicating a truncated response. While not a hard API error, this scenario implies the model couldn't complete its full answer, which significantly impacts user satisfaction. By explicitly tracking this, the system gains deeper insight into potential quality issues that wouldn't be caught by just monitoring API error rates. For unrecoverable problems like APIStatusError, the code logs the specific status and message, then re-raises the error immediately, preventing unnecessary retries for permanent issues.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a metrics collector that wraps any OpenAI chat completion call and emits a structured log line containing latency_ms, prompt_tokens, completion_tokens, finish_reason, and a boolean soft_error flag. Then call it three times with different prompts and print a summary showing average latency and the count of soft errors.

python
# openai>=1.0.0
import time, os
from openai import OpenAI

client = OpenAI()

def call_and_collect(prompt: str, model: str = "gpt-4o-mini") -> dict:
    # TODO: record start time
    # TODO: call client.chat.completions.create
    # TODO: compute latency_ms
    # TODO: extract usage.prompt_tokens, usage.completion_tokens
    # TODO: extract finish_reason from choices[0]
    # TODO: set soft_error = True if finish_reason == "length"
    # TODO: return a dict with all six fields
    pass

prompts = [
    "What is 2 + 2?",
    "List all prime numbers up to 100.",
    "Explain quantum entanglement in one sentence.",
]

results = []
for p in prompts:
    result = call_and_collect(p)
    results.append(result)
    print(result)

# TODO: print average latency_ms and count of soft_error == True across results

Quick check

  1. Your average LLM latency looks stable, but users are complaining about slow responses. Which metric would most likely reveal the problem?

  2. A call returns HTTP 200 with finish_reason 'length'. What type of failure is this, and what does it mean?

  3. You want high-volume satisfaction signals with no UI changes required. Which approach is most practical?

Self-check: Describe how you would instrument a new AI feature from scratch: which four metric categories you'd capture, how you'd tag them for slicing, and one example of a metric pair that could move in opposite directions when you make a model change.