Your LLM App Returned 200 OK. It Was Still Wrong.

Marina Petzel, Datadog18:01 · Oct 2026 · 8,160 views
Thumbnail for Your LLM App Returned 200 OK. It Was Still Wrong. Watch on YouTube
TL;DR
  1. 1

    Golden signals still show whether a GenAI application is running, but they do not show whether its output is useful or correct.

  2. 2

    GenAI cost monitoring must track token creep, model changes, uncached calls, and tags across features, users, models, and environments.

  3. 3

    Safety and quality monitoring must cover prompt injection, PII leakage, toxicity, jailbreaks, hallucinations, relevance, satisfaction, completeness, and RAG retrieval quality.

Summary

Marina Petzel explains why latency, errors, traffic, and saturation cannot fully describe the health of a generative AI application. GenAI systems can return different answers to the same prompt, have variable token-based costs, expose new attack surfaces, and produce responses whose quality falls on a spectrum. She recommends adding three monitoring layers to the traditional golden signals: cost, safety, and quality. Cost monitoring should catch expanded context windows, model changes, and repeated uncached calls, then attribute spending through tags. Safety monitoring should measure prompt injection, PII in outputs, toxic content, and jailbreak attempts. Quality monitoring should measure hallucination rate, relevance, user satisfaction, answer completeness, and retrieval quality for RAG systems. Datadog's Agent Observability product is presented as one way to monitor these signals.

Key ideas
00:01

Golden signals cannot tell you whether a GenAI answer is good

Latency, errors, traffic, and saturation still show whether an application is up and running. They do not show how well it is running. A traditional application usually gives the same output for the same input, while a generative AI system can produce different responses for the same prompt. That makes standard regression testing insufficient on its own. Petzel says teams must evaluate quality in the live environment. A server can return 200 OK while the answer is irrelevant, incomplete, inaccurate, or unhelpful to the user.

01:03

GenAI applications add four monitoring problems

Petzel describes four differences from traditional applications. Outputs are variable, so the same prompt can produce different responses. Costs are dynamic and depend on tokens, the selected model, and the context window. GenAI applications introduce attack vectors such as prompt injection, jailbreaks, and PII appearing in outputs. Their quality is subjective because an answer can be judged for relevance, accuracy, and completeness rather than as simply working or failing. These conditions require monitoring beyond infrastructure health.

03:51

Token creep, model drift, and uncached calls can drive unexpected cost

Petzel identifies three major cost problems. Token creep happens when a team expands the context window to improve quality without a financial review. She gives an example of a context window growing from 4,000 to 132,000 tokens, which she says can produce an eight-times cost increase. Model drift happens when a team moves from a cheaper, faster model to a more expensive one. Uncached calls repeat the same expensive completion when an effective caching layer is missing. She says Datadog's research found that 70% of spend can sometimes be redundant.

06:55

Cost attribution starts with mandatory tags

Petzel recommends tagging every interaction across four levels. Feature tags connect spending to functions such as chat or summarization and help product teams see which features consume the budget. User tags, such as user ID or organization name, support chargebacks, billing, and detection of abusive use. Model tags identify the model and provider so teams can compare expenses. Endpoint tags connect costs to regions or environments such as production and staging, which helps with infrastructure planning.

09:15

Safety monitoring needs explicit detection and blocking targets

The safety layer measures risks that ordinary application metrics cannot see. Prompt injection rate tracks attempts to manipulate user prompts, with pattern matching or classifier models used for detection. PII detection scans outputs for data such as social security numbers, credit cards, and other personal information. Petzel says the acceptable target for PII in production outputs is 0%. Content moderation scores estimate whether output is toxic, harmful, or biased. Jailbreak monitoring looks for attempts to bypass guardrails, with an ideal target of blocking 100% of attempts.

12:09

Quality monitoring checks whether the answer helps the user

Petzel recommends five quality metrics. Hallucination rate measures claims unsupported by grounded data and should use manual review alongside automated fact checks. Relevance scores assess whether the answer addresses the question, using methods such as BERT or embedding similarity. User satisfaction can come from thumbs-up and thumbs-down controls, star ratings, or NPS, with a suggested target of more than 80 to 85% positive feedback. Answer completeness can be evaluated by an LLM judge. RAG applications should also track retrieval quality with measures such as top-K accuracy or normalized discounted gain.

15:50

Healthy GenAI monitoring adds cost, safety, and quality to the golden signals

Petzel does not recommend removing latency, errors, traffic, and saturation from the monitoring stack. She recommends placing three additional layers on top: cost, safety, and quality. This gives teams visibility into whether the application is affordable, secure, and producing useful answers. Datadog's Agent Observability product is offered as a way to inspect these signals and identify problems in an AI application.

"We always need to look into our metrics, latency, error, traffic and saturation, but also on top of that, we need to put another layer."15:50
Who should watch
  • You operate a generative AI application and currently rely on latency, errors, traffic, and saturation to judge whether it is healthy.
  • Your LLM bill changes unexpectedly and you need a way to connect spending to features, users, models, and environments.
  • You need practical signals for prompt attacks, sensitive data leakage, hallucinations, user satisfaction, or RAG retrieval quality.