# The Transcript Looked Fine. The Call Wasn't.

Fuad Ali, Arize AI | AI Engineer | 18:00

Source: https://www.youtube.com/watch?v=42Bz0TUfeOQ
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/the-transcript-looked-fine-the-call-wasnt
Published: 2026-10-05
Tags: evals, observability, tracing, voice

## TL;DR
- A transcript can hide dead air, interruptions, misheard words, robotic tone, and an incorrect tool action in a voice call.
- Voice debugging needs audio, transcripts, traces, session-level context, and metrics in one view, with OpenInference providing a shared schema across providers.
- Audio-native evals and replayed failed traces let teams find problems, test fixes, and verify whether voice agents improve.

## Summary
Fuad Ali argues that text logs conceal important voice-agent failures. In his refund example, the transcript appears to show an agent correcting order 40 to order 14, while the audio reveals 2.4 seconds of dead air, the agent speaking over the caller, a misheard order number, and a refund processed for the wrong order. He recommends a session view that combines audio, transcript, trace data, tool calls, latency, interruptions, and eval results. OpenInference semantic conventions can normalize real-time audio events across providers such as OpenAI and Google. Ali then describes evals that operate on audio for sentiment, tone, latency, interruptions, transcription drift, and task success. His proposed workflow is observe, evaluate, improve. Investigation agents can inspect failed spans, suggest fixes, deploy them in development, and replay failed traces before an engineer approves a change.

## Key ideas
### Text logs hide the failures that shape a voice interaction
[01:07](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=67s)
Ali says voice agents are growing quickly while remaining fragile to debug. Latency, turn-taking, and transcription drift can change what a caller experiences without appearing clearly in a transcript. In the refund example, the text says the user requested order 14, the agent started a refund for order 40, and the user corrected it. The audio shows a different story: there were 2.4 seconds of dead air, the agent talked over the caller, the number 14 was misheard, and the refund went to order 40. The transcript also misses the agent's flat, robotic tone. Ali's point is direct: reading language-model output alone does not show what happened in the call.

### A useful debugging view combines the whole session
[03:56](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=236s)
Ali recommends putting audio, transcript, and trace data in the same session view. Engineers should be able to follow multiple turns, play audio inline, inspect inputs and outputs, and see time to first audio, token costs, interruption events, sentiment, and other measurements beside the interaction. The session ID links each turn into a timeline, including websocket events and whether the user or assistant was speaking. This lets an engineer hear pauses, cutoffs, and overlaps while also checking the tool calls made during the conversation. The same view can place eval results next to the spans and the audio that produced them.

### OpenInference gives different voice providers one queryable schema
[05:41](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=341s)
Ali describes audio semantic conventions mapped onto OpenInference, an open-source set of conventions for generative AI that follows OpenTelemetry conventions. The audio data can cover the session lifecycle, audio input, transcripts, conversation items, outputs, response lifecycle, and token counts. The conventions normalize providers such as OpenAI Realtime and Google Gemini Live into a schema that teams can query. That avoids bespoke parsing for each vendor and makes it easier to search large volumes of traces, filter failures, and change providers without rebuilding the instrumentation. Ali says the instrumenter can be sent to a Grafana backend or another backend.

### Replaying the full call connects audio failures to tool actions
[07:40](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=460s)
A session ID can stitch every turn and websocket event into one timeline that engineers can replay from the trace. The replay shows who spoke, when the agent paused, where either side was cut off, and which tools the agent called. In the demo, an agent retrieves a weather event, and the tool call appears inline with the surrounding audio and trace tree. Ali then returns to the refund call and identifies the three failures: 2.4 seconds before the first audio, the agent speaking over the caller, and a tool call that processed a refund for order 40 instead of order 14. Each problem would be difficult to establish from the text log alone.

### Voice evals need to inspect audio itself
[09:54](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=594s)
Ali says transcript-based judging cannot reliably capture user emotion or vocal tone. Audio evals can classify sentiment and tone from the recording, measure time to first audio against a service-level target, and flag interruptions, inaccuracy, transcription drift, and task failure. He also describes more involved agent-as-judge workflows that can inspect tools and third-party systems to check whether the agent took the required action. The eval output should sit on the relevant spans, rather than in a separate system, so teams can filter interactions by scores and inspect the exact audio and trace associated with a failure.

### Skills can create and deploy audio evals from natural-language instructions
[11:49](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=709s)
Ali gives an example instruction to create a sentiment eval with GPT audio that catches frustrated user tone on audio spans. The system can query spans, filter for audio data, create the eval, deploy it, and apply it to traces and spans. He says teams could build this themselves, but querying very large span volumes and deploying live evals as interactions arrive are difficult operational tasks. Attaching evals and monitors to spans also enables automated investigation. An investigation agent can inspect failing tool calls, find unsupported product cases, and suggest that the agent needs a new tool or a code change.

### The improvement loop runs from observation to verified fixes
[13:48](https://www.youtube.com/watch?v=42Bz0TUfeOQ&t=828s)
Ali frames voice-agent development as an observe, evaluate, improve loop. First, teams collect the interaction data. Next, they attach scores to what they observed. Then they make a fix manually or let an agent propose one, and verify the change against the original failure. His agent-experiment example describes an SRE agent that studies traces, hypothesizes that oversized tool results caused high latency, adds truncation in development, replays failed traces against the fixed endpoint, and reports whether latency decreased. An engineer then reviews a pull request with the problem description, change, and report card before approving or denying it.

## Notable quotes
- "If you're just looking at transcripts, I'll play a little audio file for you guys in a second. It doesn't really capture what's actually going on under the hood." (01:52)
- "So, what you need to do is you need to see the conversation beyond just the transcript." (03:52)
- "You want to be able to run these evals directly against the audio with evals that are based on audio models." (10:38)
- "And the loop really follows observe, evaluate, and improve." (13:48)
- "So, start seeing your agents. Add the Open Inference Instrumenter." (16:04)

## Tools & references mentioned
- Arize AI
- Arize AX
- OpenInference
- OpenTelemetry
- Grafana
- OpenAI Realtime
- Google Gemini Live
- GPT audio
- McDonald's
- White Castle
- Bojangles
- Claude
- Robinhood

## Who should watch
- You are debugging a voice agent from transcripts and need to find pauses, overlaps, mishearing, or incorrect tool calls.
- Your team runs voice systems across multiple model providers and needs shared tracing and searchable session data.
- You want to evaluate tone, latency, interruptions, and task success, then replay failed calls against proposed fixes.

## Editor's note

Fuad Ali shows how a voice-agent transcript can conceal a wrong refund, because the audio reveals a misheard order number and a tool call for order 40 instead of order 14. Kitaru records each model call and tool result in a session, then re-runs the agent code against those recorded inputs and responses. A team can test a change against the same failed run without touching real systems.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [Voice Agents: the good, the bad, and the ugly](https://aietalks.com/talks/voice-agents-the-good-the-bad-and-the-ugly) (Eddie Seagull, Fractional AI, 18:48)
- [I Monitored Crime Audio. Voice Agents Scare Me More.](https://aietalks.com/talks/i-monitored-crime-audio-voice-agents-scare-me-more) (Sumanyu Sharma, Hamming AI, 16:04)
- [From Vibes to Production: Evaluating and Shipping AI Agents That Work 201](https://aietalks.com/talks/from-vibes-to-production-evaluating-and-shipping-ai-agents-that-work) (Laurie Voss, Arize AI, 42:17)
- [Beyond Transcription: Building Voice AI That Understands Conversations](https://aietalks.com/talks/beyond-transcription-building-voice-ai-that-understands-conversations) (Hervé Bredin, pyannoteAI, 25:20)
- [Voice Agent Engineering](https://aietalks.com/talks/voice-agent-engineering) (Nik Caryotakis, SuperDial, 19:07)
