Shipping complex AI applications

Thumbnail for Shipping complex AI applications Watch on YouTube
TL;DR
  1. 1

    Production AI systems need operational workflows for tracing, evaluation, versioning, and remediation.

  2. 2

    Breaking a support triage agent into explicit stages and tool calls makes failures easier to locate and fix.

  3. 3

    A continuous loop of offline tests, production scoring, failure analysis, and prompt changes helps teams ship safely.

Summary

Giran Moodley, Mayank Soni, and Oussama Hafferssas walk through a support triage agent built for the workshop. They start with a single LLM call, then add deterministic local tools and split the workflow into stages for context collection, triage, policy review, reply writing, escalation, and finalization. Braintrust is used to trace nested calls, inspect tokens, cost, latency, metadata, and execution timelines. The workshop then adds a golden data set with deterministic scores and an LLM judge, deploys prompts, tools, and parameters in managed mode, and applies scores to production logs. The final exercise takes a failure case where the model treats an invoice request as non-urgent, tightens the prompt, reruns the evaluation, and compares the result. The argument is practical: a working prototype is only the starting point. Teams need a repeatable feedback loop that connects real production behavior to tests and code or prompt changes.

Key ideas
05:29

Production failures come from weak operating practices as often as from model behavior

Giran Moodley says sophisticated models are available, but operational rigor has not kept up with the demands of AI systems. A demo can work for two or three examples and then fail in production. Logs show what happened, while observability helps engineers understand the system's behavior across nested calls. Prompt patching can solve one issue and leave the next failure untouched. The workshop focuses on tracking changes, identifying failure modes, and creating a repeatable process for improving the application. Moodley also compares explicit AI stages with breaking traditional software into smaller services, because smaller responsibilities make problems easier to locate.

18:06

Agentic systems sit between deterministic software and non-deterministic model behavior

Oussama Hafferssas describes Trainline's systems as a mix of conventional software and model-driven behavior. Trainline handles ticket sales through a platform that works across train carriers, and its travel assistant can manage refunds, train changes, alternative journeys, and handover to human support. The team treats classic software quality checks and machine learning evaluations as complementary. Offline evaluation happens before release, while online evaluation uses production data to assess whether predictions or agent behavior remain correct. This framing matters because an agent can contain deterministic tool calls alongside model decisions, so testing only one side leaves gaps.

28:35

Start with a simple prompt, then add tools and explicit stages

The workshop's fictitious support triage application begins with one system prompt, one user input, and one structured model output. Moodley says this is a reasonable proof-of-concept starting point, but it will miss organizational nuance and edge cases. The next version adds three local tools for relevant help desk articles and account information. These tools are deterministic for the workshop, although real applications might call vector search, MCP, or external systems. The workflow is then split into stages for collecting context, triage, policy review, customer-facing reply writing, internal output, escalation, and finalization. More components create more possible failures, which makes tracing necessary.

46:19

Nested traces expose where an agent spends time and fails

Braintrust tracing records the full execution path instead of treating an agent interaction as one opaque call. The trace can contain a parent span for the ticket and child spans for functions, tools, and model calls. Moodley points to prompt and completion tokens, cost, latency, time to first token, inputs, outputs, and metadata as useful fields. The UI also provides a timeline view that resembles a waterfall, so an engineer can see which stage is slow. He says the nesting must be correct, because separate unconnected interactions make it difficult to understand the complete behavior of a multi-step workflow.

56:43

Golden data sets provide a starting point when there is no production history

For a new application, the team may not know what good output looks like. The workshop creates a golden data set with ten support cases, expected inputs, categories, and metadata. The examples cover conditions such as blocking issues, escalation policy, service-level requirements, and the structure of the response. Deterministic scoring functions check fields such as category and escalation reason. An LLM judge handles qualities that are harder to express with exact rules, including tone and helpfulness. Moodley presents this as a way to replace release decisions based on intuition with a concrete baseline that can reveal regressions when the system changes.

01:05:07

Managed prompts and parameters let teams collaborate without hiding the audit trail

The workshop moves prompts, tools, scoring functions, and parameters from the local machine into Braintrust's managed environment. A runtime mode determines whether the application uses local assets or the managed versions. In the UI, a user can change a model parameter, add a comment, save a new version, and run the application without editing source code. This gives product managers and subject matter experts a shared place to work with prompts. Moodley is careful about the limit: managed configuration does not replace version control or other access and change controls. Teams still need centralized records and automation to keep systems synchronized.

01:13:58

Online scoring should balance coverage with the cost of model judges

After offline experiments, the team applies scoring functions to live production logs through automations. Moodley recommends a higher sampling rate at the start of the monitoring process, so the team can establish a baseline and find unexpected cases. LLM-based scores can become expensive, especially when they use more capable models, so the sampling rate can later be reduced to 5 to 10 percent. Deterministic scores are cheap enough to run continuously. When there is no ground truth, his suggested path is to treat useful cases as edge cases, add them to a data set, and replay them through the evaluation environment.

01:19:13

Every production failure should become a regression case before the prompt is changed

The remediation exercise uses a support request that says an invoice export is not urgent, while the business context suggests it needs immediate attention before a board meeting. The team replays the failure, evaluates the existing behavior, tightens the prompt, and runs the updated version across the test cases. Comparing experiments shows whether the change improves the target case and whether it damages another case. Moodley describes this as a flywheel: collect information, identify a failure, remediate it, deploy the change, monitor production, and repeat. The process is more useful when it tests the full data set rather than one hand-picked example.

"It's not the prototype. It's getting to a state where we're knowing exactly what's changed in the system, how do we interact with that, and then how we systematically put a set of rigor so that we can get better and better."08:07
Who should watch
  • You have an AI prototype that works in demos, but you cannot explain regressions or production failures.
  • Your agent uses several model calls and tools, and you need traces that connect the whole execution path.
  • Product or operations colleagues need to change prompts and inspect quality without making every change through an engineer.