I Monitored Crime Audio. Voice Agents Scare Me More.

Sumanyu Sharma, Hamming AI16:04 · Sept 2026 · 3,429 views
Thumbnail for I Monitored Crime Audio. Voice Agents Scare Me More. Watch on YouTube
TL;DR
  1. 1

    Voice agents are moving into production while their confident, natural responses can still contain serious errors.

  2. 2

    A centralized prompt or architecture change can spread one failure across millions of users, unlike most local crime incidents.

  3. 3

    Reliable agents need a continuous loop of manual listening, cross-conversation analysis, adversarial testing, fixes, regression checks, and production monitoring.

Summary

Sumanyu Sharma compares his past work listening to police radio at Citizen with his current concern about voice agents. Crime usually affects a limited local group, while a voice agent can apply one bad prompt or architecture change to many users. The failures are often ordinary: skipped verification, an unauthorized discount, incorrect information, or an appointment that was never booked. Their cost depends on the situation, from annoying repetition to an unsafe food order or a failed credit-card freeze. Sharma proposes a debugging loop: identify problems, rank them by frequency and severity, make a change, test it against varied inputs, check for regressions, and keep monitoring. Manual call review gives teams useful intuition, but cross-call analysis is needed to find patterns. He also reports that Hamming's adversarial testing can break roughly one in five agents, including bypassing verification and extracting data.

Key ideas
01:29

Voice agents are leaving demos while their long tail remains unreliable

Sharma says voice products and their underlying speech and orchestration layers have improved quickly. Teams can now build experiences that are about 60% good in a short time, connect agents to calendars, CRMs, HR systems, and reservation systems, and let them take actions. The remaining long tail still causes trouble. Speech models, voice-to-voice systems, and cascaded architectures may improve the experience, but reliability remains the main problem holding back deployment at scale.

02:54

A natural voice can deliver wrong information with confidence

Sharma describes a trade-in interaction where the user became confused and a human had to call back to repair the loss of trust. He says voices sound confident and natural even when the information is wrong. His own appointment exposed the cost of this failure: he believed a physician's appointment had been booked, arrived, was not on the schedule, and lost two hours. The same error involving a grandparent or a medical procedure could have much higher consequences.

04:32

Centralized voice systems give failures a large blast radius

Sharma contrasts local crime with voice-agent deployment. Crime incidents usually affect a finite set of people connected to one situation. Voice agents are centralized, so one prompt or architecture change can affect millions of users. He estimates that at least a trillion calls happen each year and says a one percent error rate would mean ten billion incidents. Hamming monitors 10,000 agents, where he says the practical error rate is closer to ten percent.

05:23

Severity must be judged alongside frequency

The same failure mode can have very different costs. Repetition or misunderstanding is annoying, especially when it happens systematically. A voice agent that mishandles a drive-through order for someone with a peanut allergy creates a safety risk. In financial services, failing to freeze a user's credit card is a high-impact failure. Sharma's ranking puts systematic, high-impact problems first, while still treating recurring low-impact problems as worth fixing.

07:04

Reliability work is a loop from diagnosis through monitoring

Sharma lays out a process with five stages: identify problems, prioritize them by frequency and severity, understand and execute a fix, check whether the change worked without causing regressions, and continue monitoring in production. Teams should track known issues such as turnover, latency, interruptions, and automatic speech recognition problems. They also need to look for emerging behavior that only becomes visible across many conversations.

07:57

Manual call review gives intuition, while cross-call analysis finds patterns

Sharma says teams should begin by listening to individual calls rather than skipping straight to automation. That work provides detail and intuition about greetings, closing, validation, and core logic. It does not scale, so teams often move to spreadsheets and evaluation products. Those tools can score known rubrics, but they may miss new behavior spread across conversations. Sharma says Hamming spends substantial effort on cross-conversation analysis because patterns across calls reveal problems that a single-call review cannot.

10:16

A fix needs varied tests and sometimes live experimentation

Replaying one failed conversation several times is a weak test. Sharma recommends keeping the same intent while changing wording, accents, styles, and the number of intents in the interaction. This gives better coverage for deciding whether a change is actually positive. Some hypotheses, such as the opening seconds of an outbound call, are difficult to test synthetically. In those cases, he says A/B testing in real interactions is necessary.

12:07

Adversarial callers increase the risk as agents gain access

Sharma says agents are being given more data and more tools while being deployed quickly, which increases their attack surface. More natural voices can also make people easier to trick. Hamming's red-teaming product has tested agents in financial services, healthcare, and consumer settings. Sharma says the team has bypassed verification, prompted agents into revealing data, and can break about one in five agents. He recommends pre-deployment testing, broad monitoring, and 24/7 red teaming when bad interactions are costly.

"A single prompt change or an architecture change can have pretty massive implications downstream for all of the millions of users that are in the crossfire."06:37
Who should watch
  • You are deploying a voice agent that can change records, book appointments, handle payments, or access sensitive data.
  • Your team relies on call sampling or fixed evaluation rubrics and needs to find failures that appear only across many conversations.
  • You want to test whether an agent resists prompt attacks, verification bypasses, and callers trying to extract data.