The Transcript Looked Fine. The Call Wasn't.

Fuad Ali, Arize AI18:00 · Oct 2026 · 2,281 views
Thumbnail for The Transcript Looked Fine. The Call Wasn't. Watch on YouTube
TL;DR
  1. 1

    A transcript can hide dead air, interruptions, misheard words, robotic tone, and an incorrect tool action in a voice call.

  2. 2

    Voice debugging needs audio, transcripts, traces, session-level context, and metrics in one view, with OpenInference providing a shared schema across providers.

  3. 3

    Audio-native evals and replayed failed traces let teams find problems, test fixes, and verify whether voice agents improve.

Summary

Fuad Ali argues that text logs conceal important voice-agent failures. In his refund example, the transcript appears to show an agent correcting order 40 to order 14, while the audio reveals 2.4 seconds of dead air, the agent speaking over the caller, a misheard order number, and a refund processed for the wrong order. He recommends a session view that combines audio, transcript, trace data, tool calls, latency, interruptions, and eval results. OpenInference semantic conventions can normalize real-time audio events across providers such as OpenAI and Google. Ali then describes evals that operate on audio for sentiment, tone, latency, interruptions, transcription drift, and task success. His proposed workflow is observe, evaluate, improve. Investigation agents can inspect failed spans, suggest fixes, deploy them in development, and replay failed traces before an engineer approves a change.

Key ideas
01:07

Text logs hide the failures that shape a voice interaction

Ali says voice agents are growing quickly while remaining fragile to debug. Latency, turn-taking, and transcription drift can change what a caller experiences without appearing clearly in a transcript. In the refund example, the text says the user requested order 14, the agent started a refund for order 40, and the user corrected it. The audio shows a different story: there were 2.4 seconds of dead air, the agent talked over the caller, the number 14 was misheard, and the refund went to order 40. The transcript also misses the agent's flat, robotic tone. Ali's point is direct: reading language-model output alone does not show what happened in the call.

03:56

A useful debugging view combines the whole session

Ali recommends putting audio, transcript, and trace data in the same session view. Engineers should be able to follow multiple turns, play audio inline, inspect inputs and outputs, and see time to first audio, token costs, interruption events, sentiment, and other measurements beside the interaction. The session ID links each turn into a timeline, including websocket events and whether the user or assistant was speaking. This lets an engineer hear pauses, cutoffs, and overlaps while also checking the tool calls made during the conversation. The same view can place eval results next to the spans and the audio that produced them.

05:41

OpenInference gives different voice providers one queryable schema

Ali describes audio semantic conventions mapped onto OpenInference, an open-source set of conventions for generative AI that follows OpenTelemetry conventions. The audio data can cover the session lifecycle, audio input, transcripts, conversation items, outputs, response lifecycle, and token counts. The conventions normalize providers such as OpenAI Realtime and Google Gemini Live into a schema that teams can query. That avoids bespoke parsing for each vendor and makes it easier to search large volumes of traces, filter failures, and change providers without rebuilding the instrumentation. Ali says the instrumenter can be sent to a Grafana backend or another backend.

07:40

Replaying the full call connects audio failures to tool actions

A session ID can stitch every turn and websocket event into one timeline that engineers can replay from the trace. The replay shows who spoke, when the agent paused, where either side was cut off, and which tools the agent called. In the demo, an agent retrieves a weather event, and the tool call appears inline with the surrounding audio and trace tree. Ali then returns to the refund call and identifies the three failures: 2.4 seconds before the first audio, the agent speaking over the caller, and a tool call that processed a refund for order 40 instead of order 14. Each problem would be difficult to establish from the text log alone.

09:54

Voice evals need to inspect audio itself

Ali says transcript-based judging cannot reliably capture user emotion or vocal tone. Audio evals can classify sentiment and tone from the recording, measure time to first audio against a service-level target, and flag interruptions, inaccuracy, transcription drift, and task failure. He also describes more involved agent-as-judge workflows that can inspect tools and third-party systems to check whether the agent took the required action. The eval output should sit on the relevant spans, rather than in a separate system, so teams can filter interactions by scores and inspect the exact audio and trace associated with a failure.

11:49

Skills can create and deploy audio evals from natural-language instructions

Ali gives an example instruction to create a sentiment eval with GPT audio that catches frustrated user tone on audio spans. The system can query spans, filter for audio data, create the eval, deploy it, and apply it to traces and spans. He says teams could build this themselves, but querying very large span volumes and deploying live evals as interactions arrive are difficult operational tasks. Attaching evals and monitors to spans also enables automated investigation. An investigation agent can inspect failing tool calls, find unsupported product cases, and suggest that the agent needs a new tool or a code change.

13:48

The improvement loop runs from observation to verified fixes

Ali frames voice-agent development as an observe, evaluate, improve loop. First, teams collect the interaction data. Next, they attach scores to what they observed. Then they make a fix manually or let an agent propose one, and verify the change against the original failure. His agent-experiment example describes an SRE agent that studies traces, hypothesizes that oversized tool results caused high latency, adds truncation in development, replays failed traces against the fixed endpoint, and reports whether latency decreased. An engineer then reviews a pull request with the problem description, change, and report card before approving or denying it.

"If you're just looking at transcripts, I'll play a little audio file for you guys in a second. It doesn't really capture what's actually going on under the hood."01:52
Who should watch
  • You are debugging a voice agent from transcripts and need to find pauses, overlaps, mishearing, or incorrect tool calls.
  • Your team runs voice systems across multiple model providers and needs shared tracing and searchable session data.
  • You want to evaluate tone, latency, interruptions, and task success, then replay failed calls against proposed fixes.