The four metric categories in this lesson map to four different failure modes. Latency failures hurt perception and SLA compliance. Token usage failures hurt your wallet and context window constraints. Error rate failures hurt reliability. Satisfaction failures hurt retention. Each requires a different instrumentation strategy and a different alert threshold.
Latency: measure at the right granularity
Average latency is nearly useless for AI systems. A single slow request in a batch can double your mean without touching the median. What you want is p50 (the typical request), p90 (the experience for the top 10% most affected users), and p99 (the worst 1%). For interactive chat, p99 above 5-6 seconds is a hard UX problem. For background summarization pipelines, p99 of 30 seconds may be fine. The threshold depends on your product contract, not a universal standard.
Break latency down further: time-to-first-token (TTFT) and tokens-per-second (TPS). TTFT measures how long before the user sees anything, which dominates perceived responsiveness in streaming UIs. TPS measures throughput once generation starts. A high TTFT with high TPS suggests queuing or cold-start issues at the provider. A low TTFT with low TPS suggests the model is generating slowly, which might mean you're on an overloaded tier. Tag latency by model, prompt template version, and endpoint so you can isolate regressions when you ship changes.
Token usage: cost and context window are two separate concerns
You already know tokens drive cost from the tokenization lesson. The production angle is that token usage needs to be sliced by feature, user cohort, and prompt version, not just totaled per day. If your /summarize endpoint averages 1,200 input tokens per call and your /chat endpoint averages 300, a 20% traffic shift toward summarization changes your cost structure significantly. Without per-feature tracking, you'll see the cost spike and have no idea where it came from.
The second concern is context window utilization. If your average request is consuming 85% of the context window, you're one slightly longer document away from a truncation error. Track the ratio of tokens used to the model's limit per call. Alerting when p90 utilization exceeds 80% gives you runway to fix chunking or routing logic before users hit errors.
Error rates: hard vs. soft failures
Hard errors are easy: HTTP 4xx/5xx, timeouts, JSON parse failures when you expected structured output, and finish_reason == "length" (the model ran out of tokens mid-response). These are straightforward to count and alert on. A hard error rate above 0.5% in a production LLM integration is worth investigating immediately.
Soft errors are harder and more important. A soft error is when the model returns a valid HTTP 200 with plausible-looking text, but the response is wrong. This includes hallucinated facts, format violations that slip past your parser, refusals that weren't triggered by a safety violation but by a vague prompt, and responses that are technically correct but miss the user's intent. Detecting soft errors requires either an LLM-as-judge layer (covered in aidev-evaluation-llm-judge) or heuristic classifiers: does the response contain a JSON object? Does it mention the entity the user asked about? Is the response longer than five words? These weak checks catch the worst cases cheaply.
Also track finish_reason distribution. In a healthy system, the vast majority of completions should finish with "stop". A rising share of "length" finishes means your prompts or user inputs are growing, and you're silently truncating responses. Users experiencing truncated responses rarely complain, they just leave.
User satisfaction: explicit and implicit
Thumb ratings and CSAT surveys have high value but low volume. Most users won't rate a response unless it was notably good or notably bad. That selection bias skews your scores. Treat explicit ratings as a high-signal sample, not a representative measure.
Implicit signals fill in the gaps. Retry rate (did the user immediately resend a similar query?) is a strong negative signal. Copy-paste detection (did the user copy the response to clipboard?) is a positive engagement signal. Follow-up clarification rate (did the user ask a follow-up question to fix the response?) often indicates the first response was unsatisfactory. Session abandonment after an AI response is a soft negative signal. None of these are perfect, but together they give you a satisfaction proxy that correlates well with explicit ratings and has 100x the volume.
Tradeoffs and scale
At 10 users, log everything to a file and query it manually. At 10,000 users, you need a structured logging pipeline (covered in aidev-monitoring-logging) feeding a dashboard. At 10 million users, sampling becomes necessary. You cannot afford to run an LLM judge on every response at scale, so you run it on a 1-5% sample, stratified by endpoint and user segment. The key is that your sampling strategy should be deterministic and reproducible so you can rerun analysis on the same slice. Latency and token usage can be logged 100% of the time cheaply because they're just numbers. LLM-based quality evaluation is expensive and should be sampled.
Key Takeaways
- Track p50, p90, and p99 latency separately — averages hide the tail behavior that kills UX.
- Tie token usage to individual features and users, not just aggregate totals.
- Distinguish hard errors (HTTP 500, timeout) from soft errors (bad format, hallucination, refusal).
- Implicit satisfaction signals like retry rate and copy-paste behavior often outperform thumbs up/down.
Pro tips
- Track context window utilization (tokens_used / model_limit) per call, not just token count. When your p90 hits 80% of the window, you have a chunking or routing problem waiting to surface as truncation errors.
- Log finish_reason on every call. A creeping increase in 'length' finishes is a silent regression — responses are being cut off, users are getting incomplete answers, and your error rate counter shows zero because HTTP 200 came back.
- Instrument TTFT (time-to-first-token) separately from total latency, especially in streaming apps. Users tolerate long total generation times much better when they see tokens start appearing quickly. If TTFT is consistently above 1.5s, investigate provider queuing or cold-start issues before blaming generation speed.
- When you A/B test prompt changes, segment your satisfaction signals by variant from day one. Post-hoc analysis that tries to correlate a deployment timestamp with a shift in thumb ratings is almost always confounded by traffic pattern changes or other concurrent experiments.
Common pitfalls
- Mistake: Alerting on average latency instead of p99. Fix: Always set SLA alerts on p95 or p99; average latency masks tail behavior that causes the most user-visible failures.
- Mistake: Tracking only total token usage per day. Fix: Tag every call with feature name and user segment so you can pinpoint which endpoint or cohort is driving cost spikes.
- Mistake: Treating all errors as hard errors and ignoring soft failures like truncated responses or format violations. Fix: Check finish_reason and run lightweight heuristics on output structure to count soft errors separately.
- Mistake: Relying solely on explicit thumb ratings for satisfaction data. Fix: Instrument implicit signals like retry rate and session abandonment alongside explicit ratings to get statistically meaningful volume.
Which satisfaction signal to prioritize
| Option | Use when | Avoid when |
|---|---|---|
| Explicit ratings (thumbs, CSAT) | You have a UI surface where ratings fit naturally and you need high-signal labeled data for model evaluation. | Your user base is small or the interaction is high-frequency; response rates will be too low to be statistically useful. |
| Implicit behavioral signals (retry rate, copy-paste, abandonment) | You need high-volume satisfaction proxies with no UI changes; works in any product surface. | You need to distinguish between 'user rephrased for clarity' and 'user was dissatisfied'; behavioral signals are ambiguous. |
| LLM-as-judge quality scoring | You need automated quality scores on sampled traffic to catch regressions after prompt or model changes. | Evaluating every request in real time; cost and latency make it impractical above small traffic volumes without sampling. |
| Task completion rate | Your AI feature has a defined goal (form filled, booking confirmed, search query resolved) you can instrument. | The task boundary is fuzzy (open-ended chat, creative writing) where 'completion' is undefined. |
Code Example
# openai>=1.0.0, requires OPENAI_API_KEY in environment
import time, os
from openai import OpenAI
client = OpenAI()
def call_with_metrics(prompt: str) -> dict:
start = time.perf_counter()
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
)
latency_ms = (time.perf_counter() - start) * 1000
usage = response.usage
return {
"latency_ms": round(latency_ms, 1),
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"total_tokens": usage.total_tokens,
"content": response.choices[0].message.content,
"finish_reason": response.choices[0].finish_reason,
}
result = call_with_metrics("Summarize the water cycle in one sentence.")
print(result)How this code works
This code’s main job is to send a request to an AI model and meticulously record important performance and cost metrics. It starts by setting up the OpenAI client, which requires an OPENAI_API_KEY to connect to the AI service. The core logic resides within the call_with_metrics function. This function uses time.perf_counter() to accurately mark the beginning of an AI interaction, ensuring precise measurement of how long the AI takes to respond before client.chat.completions.create is invoked.
Inside call_with_metrics, the actual AI call happens using client.chat.completions.create. This specifies the model to use (like gpt-4o-mini, a common choice for its balance of cost and speed) and the messages containing the prompt. Immediately after the response, time.perf_counter() is called again to calculate latency_ms, representing the time taken in milliseconds. The code then extracts detailed token usage information from response.usage, including prompt_tokens (input), completion_tokens (output), and total_tokens. Finally, it returns a neat dictionary containing all these metrics along with the AI’s generated content and the finish_reason for the response, providing a comprehensive overview of the AI interaction.
Production-grade example
Adds retries with backoff, per-feature structured logs, timeout handling, and soft-error detection.
# openai>=1.0.0, structlog>=23.0.0
import time, os, asyncio
from openai import AsyncOpenAI, APITimeoutError, RateLimitError, APIStatusError
import structlog
log = structlog.get_logger()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
MAX_RETRIES = 3
BASE_BACKOFF = 1.0 # seconds
async def call_with_full_metrics(
prompt: str,
model: str = "gpt-4o-mini",
feature: str = "unknown",
user_id: str = "anonymous",
) -> dict:
attempt = 0
while attempt < MAX_RETRIES:
start = time.perf_counter()
try:
response = await client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
)
latency_ms = round((time.perf_counter() - start) * 1000, 1)
usage = response.usage
finish_reason = response.choices[0].finish_reason
# Detect soft error: truncated response
soft_error = finish_reason == "length"
log.info(
"llm_call_success",
model=model,
feature=feature,
user_id=user_id,
latency_ms=latency_ms,
prompt_tokens=usage.prompt_tokens,
completion_tokens=usage.completion_tokens,
total_tokens=usage.total_tokens,
finish_reason=finish_reason,
soft_error=soft_error,
attempt=attempt + 1,
)
return {
"content": response.choices[0].message.content,
"latency_ms": latency_ms,
"total_tokens": usage.total_tokens,
"soft_error": soft_error,
}
except RateLimitError:
wait = BASE_BACKOFF * (2 ** attempt)
log.warning("llm_rate_limited", feature=feature, attempt=attempt + 1, retry_in=wait)
await asyncio.sleep(wait)
attempt += 1
except APITimeoutError:
latency_ms = round((time.perf_counter() - start) * 1000, 1)
log.error("llm_timeout", feature=feature, latency_ms=latency_ms, attempt=attempt + 1)
attempt += 1
except APIStatusError as e:
log.error("llm_api_error", feature=feature, status_code=e.status_code, message=str(e))
raise # non-retryable
log.error("llm_max_retries_exceeded", feature=feature, max_retries=MAX_RETRIES)
raise RuntimeError(f"LLM call failed after {MAX_RETRIES} attempts")How this code works
This code defines a robust way to interact with OpenAI's language models, specifically designed to gather critical performance data for AI observability lessons. The primary function, call_with_full_metrics, wraps an OpenAI API call, focusing on tracking key metrics like latency_ms, prompt_tokens, completion_tokens, and the finish_reason to understand model behavior. It ensures reliability by implementing a retry mechanism with exponential backoff for RateLimitError exceptions, and it logs APITimeoutError events, reflecting common real-world challenges. Structured logging with structlog records all these details, providing a rich dataset for analysis in an AI monitoring system.
Beyond simple success/failure, the code subtly captures a soft_error when the finish_reason is length, indicating a truncated response. While not a hard API error, this scenario implies the model couldn't complete its full answer, which significantly impacts user satisfaction. By explicitly tracking this, the system gains deeper insight into potential quality issues that wouldn't be caught by just monitoring API error rates. For unrecoverable problems like APIStatusError, the code logs the specific status and message, then re-raises the error immediately, preventing unnecessary retries for permanent issues.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a metrics collector that wraps any OpenAI chat completion call and emits a structured log line containing latency_ms, prompt_tokens, completion_tokens, finish_reason, and a boolean soft_error flag. Then call it three times with different prompts and print a summary showing average latency and the count of soft errors.
# openai>=1.0.0
import time, os
from openai import OpenAI
client = OpenAI()
def call_and_collect(prompt: str, model: str = "gpt-4o-mini") -> dict:
# TODO: record start time
# TODO: call client.chat.completions.create
# TODO: compute latency_ms
# TODO: extract usage.prompt_tokens, usage.completion_tokens
# TODO: extract finish_reason from choices[0]
# TODO: set soft_error = True if finish_reason == "length"
# TODO: return a dict with all six fields
pass
prompts = [
"What is 2 + 2?",
"List all prime numbers up to 100.",
"Explain quantum entanglement in one sentence.",
]
results = []
for p in prompts:
result = call_and_collect(p)
results.append(result)
print(result)
# TODO: print average latency_ms and count of soft_error == True across resultsQuick check
Your average LLM latency looks stable, but users are complaining about slow responses. Which metric would most likely reveal the problem?
A call returns HTTP 200 with finish_reason 'length'. What type of failure is this, and what does it mean?
You want high-volume satisfaction signals with no UI changes required. Which approach is most practical?