# AIE Talks > Summaries of talks from the AI Engineer YouTube channel (https://www.youtube.com/@aiDotEngineer). Each talk page has a three-bullet TL;DR, a summary, five to nine key ideas with links into the video at the second they start, quotes, and tags. Curated packs put talks in a deliberate order. Every talk page and pack page also exists as plain Markdown: add .md to any talk or pack URL, or send Accept: text/markdown to the same URL. ## Packs - [Coding agents on real codebases](https://aietalks.com/packs/coding-agents-on-real-codebases): Coding agents do fine in a fresh repo and make a mess of a ten year old one. - [Agents in production: reliability, evals and cost](https://aietalks.com/packs/agents-in-production): A demo agent has one failure mode, a wrong answer. ## Talks - [Agent Spending Without Controls](https://aietalks.com/talks/agent-spending-without-controls): Rodrigo Coelho & Pranav Maheshwari, Edge & Node. Rodrigo Coelho and Pranav Maheshwari argue that agentic commerce needs two pieces of infrastructure: a way for agents to pay for useful tools and controls... - [Beyond the Lethal Trifecta: Agentic Commerce on the Open Internet](https://aietalks.com/talks/beyond-the-lethal-trifecta-agentic-commerce-on-the-open-internet): David Levine, Kiduna Club. David Levine traces his idea of agentic commerce back to LambdaMOO, a text-based virtual world he joined in 1993. He says its sense of reality... - [Multimodal Collaborative Agents for Next-Gen Commerce](https://aietalks.com/talks/multimodal-collaborative-agents-for-next-gen-commerce): Nidhi Kaushik Vyas, Google DeepMind. Nidhi Kaushik Vyas presents a framework for commerce agents that work with incomplete, subjective intent. She argues that users often arrive with a vibe rather... - [Teaching agents to pay](https://aietalks.com/talks/teaching-agents-to-pay): Anna Spysz, Stripe. Anna Spysz builds a shopping agent to find headphones for recording, mixing, and mastering music. The example introduces the infrastructure needed for an agent to... - [The End of the Static Screen: Architecting Intent-Driven UX](https://aietalks.com/talks/the-end-of-the-static-screen-architecting-intent-driven-ux): Gus Iwanaga, commercetools. Gus Iwanaga argues that software still makes people adapt to its fixed interfaces, even though AI can generate more personalized experiences. He demonstrates why unconstrained... - [When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS](https://aietalks.com/talks/when-ai-agents-pay-and-sellers-monetize-building-x402-apps-on-aws): Anil Nadiminti, AWS. Anil Nadiminti argues that web payment systems built around human subscriptions and card transactions do not fit autonomous agents. Agents may need to pay a... - [Why Your AI Agent Needs a Wallet: USDC and Nanopayments](https://aietalks.com/talks/why-your-ai-agent-needs-a-wallet-usdc-and-nanopayments): Harshal Bhangale, Circle. Harshal Bhangale argues that payment is a practical limit on AI agents. Agents can research across many sources, but they stop at paywalls because existing... - [x402 Isn't Good (Yet)](https://aietalks.com/talks/x402-isnt-good-yet): Jan Curn, Apify. Jan Curn presents x402 as an exciting protocol that still has serious implementation problems. Apify operates a marketplace with about 45,000 tools, so it needs... - [Your Agent Just Authorized What?!](https://aietalks.com/talks/your-agent-just-authorized-what): Jay Mok & Ben Coumes, PayPal. Jay Mok and Ben Coumes present a framework for deciding how agents should be authorized. Every system needs to answer three questions: did a human... - [SOTA Generative Media Panel](https://aietalks.com/talks/sota-generative-media-panel): Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind. The panel covers Google DeepMind's recent generative media work, including Nano Banana 2 Lite and the Gemini Omni Flash APIs. Nicole Brichtova describes image-to-video workflows,... - [Agentic Sites: Building Hyper Personalized Websites](https://aietalks.com/talks/agentic-sites-building-hyper-personalized-websites): Carlos Sanchez, Adobe. Carlos Sanchez presents Adobe's approach to agentic sites, where a website adapts itself to a visitor's intent in real time. The system records browsing signals,... - [Agents Are Where Microservices Were in 2015](https://aietalks.com/talks/agents-are-where-microservices-were-in-2015): Roberto Milev & Uday Kanagala, Navan. Roberto Milev and Uday Kanagala compare the current agent wave with the microservices wave. Their advice is to learn the basic unit before adding orchestration.... - [AI Agents Are Just Distributed Systems Now](https://aietalks.com/talks/ai-agents-are-just-distributed-systems-now): Salman Munaf, TikTok. Salman Munaf argues that tool-using agents should be designed as distributed systems. A text-only model could produce a bad answer, but an agent can call... - [From Tokenmaxxing to Trusted Throughput](https://aietalks.com/talks/from-tokenmaxxing-to-trusted-throughput): Mingsheng Hong, Ironclad. Mingsheng Hong argues that engineering teams should stop treating token usage as a target. A dashboard can reveal adoption gaps, sudden usage bursts, or unusual... - [Tell the Robot What You Want](https://aietalks.com/talks/tell-the-robot-what-you-want): Sandhya Subramani, AWS. Sandhya Subramani demonstrates Scout, a four-legged robot with a Raspberry Pi, a SIM card, and a 4G connection. Scout already has preset movement policies, but... - [The Half Life of Agent Infrastructure](https://aietalks.com/talks/the-half-life-of-agent-infrastructure): Ben Kus, Box. Ben Kus argues that infrastructure advice built around long-lived systems does not fit AI agents. Databases, identity controls, storage, and engineering practices can remain useful... - [The Signal Layer: What to Build When Anything Can Be Built](https://aietalks.com/talks/the-signal-layer-what-to-build-when-anything-can-be-built): Lena Hall, Akamai. Lena Hall argues that AI has made average software and content cheap to produce, because competitors can point similarly capable systems at the same questions... - [Tribal Dungeons of Global Shipping: AI Agents at Global Scale](https://aietalks.com/talks/tribal-dungeons-of-global-shipping-ai-agents-at-global-scale): Dmitry Buykin, Maersk. Dmitry Buykin describes agent work in global shipping, where a shipment is really an orchestration of many parallel state machines. The easy cases are already... - [Which AI startups actually land enterprise contracts?](https://aietalks.com/talks/which-ai-startups-actually-land-enterprise-contracts): Brian Lewis, Millennium. Brian Lewis explains enterprise AI buying from the perspective of a buyer at Millennium, while speaking personally rather than for his employer. For one internal... - [AI Evals for Cross-Functional Teams](https://aietalks.com/talks/ai-evals-for-cross-functional-teams): Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash. DoorDash's GenAI platform team began with the assumption that evals were an engineering concern. Different product groups needed different forms of evaluation, including session-level judgments,... - [AI-Native Organisations Run on Skills: How to Structure and Scale Them](https://aietalks.com/talks/ai-native-organisations-run-on-skills-how-to-structure-and-scale-them): Imad Touil, QuantumBlack. Imad Touil argues that skills are where an AI-native organization's practical know-how becomes executable. He places them inside a larger agentic stack with workflows, hooks,... - [Building the Engine While Flying the Plane: Launching the Figma MCP Server](https://aietalks.com/talks/building-the-engine-while-flying-the-plane-launching-the-figma-mcp-server): Jesse Lumarie, Figma. Jesse Lumarie describes how Figma built and shipped its first MCP server in about three months while the protocol and client ecosystem were changing. He... - [Building uReview, Uber's Multi-Agent Code Review Engine](https://aietalks.com/talks/building-ureview-ubers-multi-agent-code-review-engine): Will Bond & Ameya Ketkar, Uber. Will Bond and Ameya Ketkar explain why Uber built uReview instead of buying an automated code review product. Uber needed support for Phabricator, the same... - [From AI-Assisted to AI-Native: Building a Frontier Development Team](https://aietalks.com/talks/from-ai-assisted-to-ai-native-building-a-frontier-development-team): Clare Liguori, AWS. Clare Liguori describes a shift from chat-based coding assistance to what Amazon calls frontier development. Frontier developers write only 1 to 2 percent of their... - [How do you diffuse AI into the real world?](https://aietalks.com/talks/how-do-you-diffuse-ai-into-the-real-world): Varun Shenoy, Long Lake. Varun Shenoy argues that the hard part of AI is no longer demonstrating what models can do. It is getting them to complete economically useful... - [How to Avoid Disaster When Vibe-Coding a Billing Engine](https://aietalks.com/talks/how-to-avoid-disaster-when-vibe-coding-a-billing-engine): Andrew Garvin, Stripe. Andrew Garvin demonstrates a Stripe Projects workflow that uses a coding agent to create a Metronome billing sandbox from a natural-language request. His example copies... - [How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage)](https://aietalks.com/talks/how-to-get-your-org-to-adopt-coding-agents-without-shipping-garbage): Eyal Blum, Figma. Eyal Blum describes AI adoption at Figma as a three-act process. Engineers first find small tasks where agents work well, then lose trust when they... - [Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons](https://aietalks.com/talks/productionizing-llm-gateways-architecture-tradeoffs-and-hard-lessons): Kanish Manuja, Twilio. Kanish Manuja explains the engineering choices behind production LLM gateways. A gateway sits between applications and model providers, handling routing, authentication, fallbacks, rate limits, and... - [Your Code Has Bugs. Lean4 Has Proofs: Formal Verification for Engineers](https://aietalks.com/talks/your-code-has-bugs-lean4-has-proofs-formal-verification-for-engineers): Varun Pant, AWS. Varun Pant argues that coding agents have made ordinary software checks insufficient. Model-based grading is probabilistic, tests cover selected inputs, and human review cannot keep... - [Can LLMs Write Fast Multi-GPU Kernels?](https://aietalks.com/talks/can-llms-write-fast-multi-gpu-kernels): Simran Arora, Together AI. Simran Arora explains why multi-GPU communication has become a larger performance problem as GPU compute has advanced faster than the links between GPUs. She introduces... - [How Anthropic Builds: Lessons from Labs](https://aietalks.com/talks/how-anthropic-builds-lessons-from-labs): Mike Krieger, Anthropic. Mike Krieger describes how his work changed after moving from Anthropic's chief product officer role into an individual contributor position at Labs. He now gives... - [How to Generate Mergeable Code with a Context Engine](https://aietalks.com/talks/how-to-generate-mergeable-code-with-a-context-engine): Peter Werry, Unblocked. Peter Werry argues that agents need more than access to a codebase or a large context window. Before agents, engineers carried organizational context by reading... - [KV Cache-Aware Routing and P/D Disaggregation on Kubernetes](https://aietalks.com/talks/kv-cache-aware-routing-and-p-d-disaggregation-on-kubernetes): Yuchen Fama & Ashish Kamra, Red Hat. Yuchen Fama and Ashish Kamra describe two ways to run inference for agentic workloads on Kubernetes. The first is KV cache-aware routing. Agent sessions can... - [The Agentic Commerce Stack](https://aietalks.com/talks/the-agentic-commerce-stack): Ahnaf Prio, Best Buy. Ahnaf Prio describes what Best Buy has learned while working on agentic commerce. Shopping already makes up about 45% of sessions on major AI assistants,... - [AI in GTM at Notion](https://aietalks.com/talks/ai-in-gtm-at-notion): Flora Liu, Notion. Flora Liu describes Notion's effort to unify self-serve growth and sales-assisted GTM into one decisioning system. Customer data previously lived across Salesforce, Gong, Outreach, ZoomInfo,... - [Building GTM AI Agents: Lessons from Deploying to 6,000 Users](https://aietalks.com/talks/building-gtm-ai-agents-lessons-from-deploying-to-6-000-users): Sait Izmit, Snowflake. Sait Izmit describes Snowflake's internal go-to-market assistant, which launched in September and now serves about 6,000 users. His approach starts with quality. Before connecting more... - [GTM Engineering: The Technical Bits](https://aietalks.com/talks/gtm-engineering-the-technical-bits): Everett Berry, Clay. Everett Berry describes GTM engineering as the technical work needed to let go-to-market teams ship data, automations, and campaigns at an engineering-like pace. He focuses... - [How AI Agents Let GTM Teams Scale](https://aietalks.com/talks/how-ai-agents-let-gtm-teams-scale): Justin Joyce, Cloudflare. Justin Joyce argues that traditional go-to-market operations break down as teams grow. Operations staff rebuild analysis in spreadsheets or maintain dashboards that cannot answer every... - [How We Got LLMs to Recommend Our Open Source Library](https://aietalks.com/talks/how-we-got-llms-to-recommend-our-open-source-library): Christopher Burns, Inth. Christopher Burns explains how Inth made its open source consent banner library, c15t, easier for language models and coding agents to understand. The work started... - [Knowledge Systems: The New GTM Stack](https://aietalks.com/talks/knowledge-systems-the-new-gtm-stack): Jeffrey Wang, Exa. Jeffrey Wang argues that go-to-market is an engineering problem because its work is mostly about collecting, organizing, and acting on data. Exa maintains a live... - [Reverse-Engineering the AI Buyer](https://aietalks.com/talks/reverse-engineering-the-ai-buyer): Aliisa Rosenthal, Acrew Capital. Aliisa Rosenthal uses OpenAI's early enterprise experience to explain how AI companies should build their go-to-market motion. OpenAI had strong demand after ChatGPT launched, but... - [The Building Blocks of GTM Orchestration](https://aietalks.com/talks/the-building-blocks-of-gtm-orchestration): Arman Vaziri, Ramp. Arman Vaziri describes GTM orchestration as the ability to describe a playbook, experiment, or campaign and distribute its execution across channels. He argues that the... - [The Death of Developer Advocates](https://aietalks.com/talks/the-death-of-developer-advocates): Stephanie Jarmak, Sourcegraph. Stephanie Jarmak argues that developer relations is changing because developers now work with agents, and agents have become users and recommenders of software tools. An... - [The Missing Layer in Agentic AI](https://aietalks.com/talks/the-missing-layer-in-agentic-ai): Giedrius Šteimantas, Oxylabs. Giedrius Šteimantas uses a personal shopping agent to explain the infrastructure layer that agentic systems need when they work on the open web. The original... - [Einstein Arena: Harnessing Collective Agent Intelligence for Open Science](https://aietalks.com/talks/einstein-arena-harnessing-collective-agent-intelligence-for-open-science): James Zou, Together AI. James Zou argues that agent systems should be designed around environments rather than fixed workflows. A workflow tells an agent which steps, tools, and prompts... - [Agent Frameworks Considered Harmful](https://aietalks.com/talks/agent-frameworks-considered-harmful): Rémi Louf, .txt. Rémi Louf describes taking two weeks away from running .txt to build a background-agent system for his own morning routine. The system processes market news,... - [FinOps for AI Agents: Who Spent All the Tokens?](https://aietalks.com/talks/finops-for-ai-agents-who-spent-all-the-tokens): Tisha Chawla & Susheem Koul, Microsoft. Tisha Chawla and Susheem Koul describe Token Ops, a control plane for managing the cost of AI agent runs. They argue that SaaS and cloud... - [Give the Agent a Budget, Not a Token](https://aietalks.com/talks/give-the-agent-a-budget-not-a-token): Sachin Malhotra, Anthropic. Sachin Malhotra argues that token scopes are too blunt for agents working in production. A token answers whether an operation is allowed, but it does... - [Preferences Over Benchmarks: Model Routing](https://aietalks.com/talks/preferences-over-benchmarks-model-routing): Archana Kamath & Tyler Gillam, DigitalOcean. Archana Kamath and Tyler Gillam argue that choosing one model by its position on a public benchmark misses the conditions that determine whether it fits... - [The Agent Behind the Curtain: Building the Oz Cloud Agent Platform](https://aietalks.com/talks/the-agent-behind-the-curtain-building-the-oz-cloud-agent-platform): Safia Abdalla, Warp. Safia Abdalla explains how Warp built a cloud platform for agents around a simple principle: the platform should absorb complexity before it reaches the user.... - [What If Your Chip Design Team Moved Like a Single Body?](https://aietalks.com/talks/what-if-your-chip-design-team-moved-like-a-single-body): Abduallah Mohamed, AIDAChip. Abduallah Mohamed argues that chip design teams lose more time to alignment than to a lack of individual engineering skill. In conversations with about 15... - [Agentic SDLC at Uber](https://aietalks.com/talks/agentic-sdlc-at-uber): Uday Kiran Medisetty & Adam Huda, Uber. Uday Kiran Medisetty describes six infrastructure pieces behind Uber's agentic software development work. A model gateway gives every request a caller, team, and project identity,... - [Building Agents Is Trivial Now, Context Is the Next Frontier](https://aietalks.com/talks/building-agents-is-trivial-now-context-is-the-next-frontier): Jeff Ng, Unblocked. Jeff Ng argues that deploying an agent has become much easier. Earlier, teams had to build checkpointing, state persistence, sandbox isolation, and observability before an... - [The Missing Layer: Design Taste in AI Agents](https://aietalks.com/talks/the-missing-layer-design-taste-in-ai-agents): Hassan El Mghari, Together AI. Hassan El Mghari builds around ten apps a year without being a designer. He says design and UX have helped some of those apps reach... - [How I automate my own job at Hugging Face using agents](https://aietalks.com/talks/how-i-automate-my-own-job-at-hugging-face-using-agents): Niels Rogge, Hugging Face. Niels Rogge describes the Community Science team at Hugging Face as a "Google Drive to the Hub" team. Its members find papers whose models or... - [IT Admin for the AI Workforce](https://aietalks.com/talks/it-admin-for-the-ai-workforce): Sarthak Aggarwal, Decawork. Sarthak Aggarwal argues that autonomous agents should be managed like software employees. Once an agent can read private context, make decisions, call tools, and act... - [Prototyping as Leadership: How a CTO Ships with AI Agents](https://aietalks.com/talks/prototyping-as-leadership-how-a-cto-ships-with-ai-agents): Hursh Agrawal, The Browser Company. Hursh Agrawal argues that leaders should keep building even when their calendars fill with meetings. He has 15 or more recurring meetings a week, seven... - [The Era of Compound Engineering](https://aietalks.com/talks/the-era-of-compound-engineering): Kieran Klaassen, Every/Cora. Kieran Klaassen describes rebuilding Cora, an AI-native email client, as a single engineer with support from specialists. His work moved through several bottlenecks: first code... - [The Last Human Code Review: Building Trust in AI-Generated Code](https://aietalks.com/talks/the-last-human-code-review-building-trust-in-ai-generated-code): Itamar Friedman, Qodo. Itamar Friedman argues that AI-generated code has moved the bottleneck from writing software to reviewing and governing it. He frames code review as having two... - [Unlock Agent Autonomy: The Runtime for AI-Native Systems](https://aietalks.com/talks/unlock-agent-autonomy-the-runtime-for-ai-native-systems): Tushar Jain, Docker. Tushar Jain argues that agent autonomy is limited by safety rather than intelligence. An agent investigating an incident may reasonably ask for logs, GitHub history,... - [Your Agent Evolved. Your Evals Didn't.](https://aietalks.com/talks/your-agent-evolved-your-evals-didnt): Ameya Bhatawdekar, Braintrust. Ameya Bhatawdekar describes five generations of AI application architecture, from a single prompt through chains, ReAct loops, workflow graphs, and newer agent systems with memory,... - [Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards](https://aietalks.com/talks/your-fine-tuned-model-is-tech-debt-a-50x-roi-house-of-cards): Dan Bjornn, Lease End. Dan Bjornn describes Lease End's LLM messaging application, which classified customer intent for questions, scheduling, reminders, and sales calls. Retrieval worked reasonably well but missed... - [200 Million Patient Interactions Later](https://aietalks.com/talks/200-million-patient-interactions-later): Vivek Muppalla, Hippocratic AI. Vivek Muppalla argues that healthcare has been shaped by scarcity. Since clinicians do not have enough time to call everyone, systems triage patients and focus... - [AI is the World's Largest Relationship Therapist](https://aietalks.com/talks/ai-is-the-worlds-largest-relationship-therapist): Clay Cockrell & Tony Fabrikant, CoupleWork AI. Clay Cockrell argues that AI has become the world's largest relationship therapist because people turn to general-purpose assistants after arguments, often late at night. He... - [Don't Be Data Poor](https://aietalks.com/talks/dont-be-data-poor): Anuj Iravane, Anterior. Anterior works with scanned medical records that arrive mostly by fax. These records contain handwriting, tables, checkboxes, images, and information spanning a patient's clinical history.... - [From Ambient Documentation to Clinical Intelligence](https://aietalks.com/talks/from-ambient-documentation-to-clinical-intelligence): Chaitanya Asawa, Abridge. Chaitanya Asawa describes Abridge's move from ambient clinical documentation toward clinical intelligence. The company started by reducing the time clinicians spend writing notes after visits.... - [Guardrails First: Engineering Member-Facing Health AI](https://aietalks.com/talks/guardrails-first-engineering-member-facing-health-ai): Rashi Agrawal, Hinge Health. Rashi Agrawal describes how Hinge Health approaches safety for member-facing healthcare AI. She starts with failures from consumer health assistants, including dangerous diet advice and... - [Healthcare's Agent Bytecode: X12 as the Harness for AI Agents](https://aietalks.com/talks/healthcares-agent-bytecode-x12-as-the-harness-for-ai-agents): Vasant Kearney, Onlay. Vasant Kearney argues that healthcare agents need a constrained execution layer because claims involve many dependent actions. A model may query a database, inspect an... - [How to build an AI-Native Health Company](https://aietalks.com/talks/how-to-build-an-ai-native-health-company): Dan Feng, Maven Clinic. Dan Feng describes Maven Clinic's move from a traditional technology company toward an AI-native model. The change covers internal work, customer products, hiring, performance reviews,... - [Shipping AI to a Million Patients Without an A/B Test](https://aietalks.com/talks/shipping-ai-to-a-million-patients-without-an-a-b-test): Jared Joselowitz, Ufonia. Jared Joselowitz explains how Ufonia evaluates Dora, a voice agent that calls patients for post-operative follow-ups and pre-operative checks. Dora has made around 200,000 clinical... - [Trading Desks to Clinical Trials: Parallels in Applied Vertical AI](https://aietalks.com/talks/trading-desks-to-clinical-trials-parallels-in-applied-vertical-ai): Ayush Bhardwaj, Allos AI. Ayush Bhardwaj compares his work at a hedge fund with his current role at Allos AI, a pharma technology company. The industries have different timelines... - [Why Your Enterprise Tech Stack Isn't Ready for AI Agents](https://aietalks.com/talks/why-your-enterprise-tech-stack-isnt-ready-for-ai-agents): Christopher Lovejoy, Anthropic & Saul Howard, Anterior. Christopher Lovejoy and Saul Howard describe the gap between an impressive enterprise AI proof of concept and a system that can enter production. Their healthcare... - [Building an Agentic Video Editor for Mass Consumer](https://aietalks.com/talks/building-an-agentic-video-editor-for-mass-consumer): Ekaterina Deyneka, Reelful. Ekaterina Deyneka presents Reelful as an agentic video editor for people who record content but rarely publish it. A user adds media and a direction,... - [Generative Video at the Speed of Light](https://aietalks.com/talks/generative-video-at-the-speed-of-light): Keegan McCallum, uRun. Keegan McCallum argues that generative video progress should be measured by efficiency and generation length as well as visual quality. His example is Helios, a... - [Infra behind Krea 2: How to train and serve at scale](https://aietalks.com/talks/infra-behind-krea-2-how-to-train-and-serve-at-scale): Gabriel Jorge Menezes, Krea.ai. Gabriel Jorge Menezes describes the infrastructure behind Krea 2, a diffusion transformer trained from scratch on thousands of GPUs. At larger cluster sizes, runs crashed... - [The Next Game Engine Won't Have a Manual](https://aietalks.com/talks/the-next-game-engine-wont-have-a-manual): Arturo Nunez, Nereu. Arturo Nunez argues that conventional game engines force developers to think in engine vocabulary before they can express game ideas. Controlling a character may require... - [The Next Medium: Why Real-Time Interactive Video Changes Everything](https://aietalks.com/talks/the-next-medium-why-real-time-interactive-video-changes-everything): Ahmed Ahres, Reactor. Ahmed Ahres argues that real-time video is a new medium rather than a faster version of generated video. Current systems return a file after a... - [Training Krea 2: What Matters in Generative Model Training](https://aietalks.com/talks/training-krea-2-what-matters-in-generative-model-training): Sangwu Lee, Krea.ai. Sangwu Lee explains how Krea trained Krea 2 and why the team focused on stylistic diversity rather than the highly consistent outputs of slower production... - [Voice agents with Realtime Video](https://aietalks.com/talks/voice-agents-with-realtime-video): Sidney Primas, LemonSlice. Sidney Primas describes LemonSlice's attempt to make video avatars indistinguishable from people on a video call. The company recently deployed a full-body Teddy Roosevelt avatar... - [While My Guitar Gently Speaks](https://aietalks.com/talks/while-my-guitar-gently-speaks): Todd Fisher, Philo Ventures. Todd Fisher describes a long-running project to make a guitar speak. He places it in the history of guitar technology, from pickups and amplifiers to... - [Context Engineering in 2026](https://aietalks.com/talks/context-engineering-in-2026): Louis-François Bouchard, Omar Solano & Samridhi Vaid, Towards AI. Louis-François Bouchard, Omar Solano, and Samridhi Vaid describe the context problems behind Towards AI's open-source AI tutor. The model has a finite context window and... - [How to Kill the Code Review](https://aietalks.com/talks/how-to-kill-the-code-review): Ankit Jain, Aviator. Ankit Jain argues that the old code review process is already breaking down. Code volume has risen, review waits have grown, and more than 30%... - [Security Firewall for Agents](https://aietalks.com/talks/security-firewall-for-agents): Ryan Dahl, Deno. Ryan Dahl describes how Deno uses agents to handle Deno Deploy incidents with read and write access to Postgres, Kubernetes, ClickHouse, AWS, GitHub, and Slack.... - [Bringing agents onto the World Wide Web](https://aietalks.com/talks/bringing-agents-onto-the-world-wide-web): Paul Klein IV, Browserbase. Paul Klein IV argues that browser agents are ready for more work than current products deliver. Recent models can handle longer tasks and use interfaces,... - [Computer Use at the Edge of the Statistical Precipice](https://aietalks.com/talks/computer-use-at-the-edge-of-the-statistical-precipice): Pierluca D'Oro, Programma Labs. Pierluca D'Oro argues that computer-use benchmarks can reward agents that never understand or inspect the screen. His replay agent records a successful trajectory for every... - [Computer-use models will agentify the web, not APIs](https://aietalks.com/talks/computer-use-models-will-agentify-the-web-not-apis): Dhruv Batra, Yutori. Dhruv Batra accepts that agents will perform much of the action on the web, but rejects the idea that the web will quickly provide APIs... - [From RL to IRL](https://aietalks.com/talks/from-rl-to-irl): Gaurav Mishra, Amazon AGI Lab. Gaurav Mishra explains why reinforcement learning produced strong coding agents and why the same approach breaks when agents operate websites and other real interfaces. Coding... - [How Web Data Infrastructure Powers the Next Generation of AI](https://aietalks.com/talks/how-web-data-infrastructure-powers-the-next-generation-of-ai): Patricija Žemaitytė, Oxylabs. Patricija Žemaitytė argues that AI systems depend on infrastructure that can collect and deliver fresh public web data. She explains this through three projects. A... - [The Dark Arts of Web Automation: Teaching Agents to Use Websites Like Humans](https://aietalks.com/talks/the-dark-arts-of-web-automation-teaching-agents-to-use-websites-like-humans): Corey Gallon, Rexmore. Corey Gallon presents a method for making browser agents work with websites that resist automation. His premise is that a browser driven through Chrome DevTools... - [The Rise of CaaS: Context-as-a-Service for Agentic AI](https://aietalks.com/talks/the-rise-of-caas-context-as-a-service-for-agentic-ai): Omer Primor, Bright Data. Omer Primor argues that agents should treat the web as an ongoing source of context rather than a static data store. Social content can become... - [Beyond Static Intelligence: Evaluating Continual Learning](https://aietalks.com/talks/beyond-static-intelligence-evaluating-continual-learning): Parth Asawa, UC Berkeley. Parth Asawa argues that standard language-model evaluation quietly assumes the system forgets everything after each task. That setup measures isolated capability, while continual learning asks... - [Bringing Continual Learning into Enterprises](https://aietalks.com/talks/bringing-continual-learning-into-enterprises): Samuel Denton, Applied Compute. Samuel Denton presents continual learning as two connected spectrums. Traces range from a one-time offline dump to a system where serving and training share one... - [Designing Agents (The Floor Is the Frontier)](https://aietalks.com/talks/designing-agents-the-floor-is-the-frontier): Ben Hylak, Raindrop. Ben Hylak argues that improving production agents means raising their floor, the worst behavior they can produce, rather than chasing benchmark scores or impressive edge-case... - [Gradient-Free Continual Learning](https://aietalks.com/talks/gradient-free-continual-learning): Sara Hooker, Adaption. Sara Hooker argues that frontier AI has been shaped by an unusually narrow path: the right PhD, the right lab, the right project, and access... - [Improving Agents is a Data Mining Problem](https://aietalks.com/talks/improving-agents-is-a-data-mining-problem): Vivek Trivedy, LangChain. Vivek Trivedy argues that improving an agent starts with operating it in the real world and saving everything it does. Tool calls, messages, API calls,... - [Intelligence + Continual Learning = Expertise](https://aietalks.com/talks/intelligence-continual-learning-expertise): Yu Su, NeoCognition. Yu Su separates intelligence from expertise. Intelligence is the ability to reason through an unfamiliar problem using the context, tools, and instructions provided. Expertise is... - [Lessons from Studying Every Memory System](https://aietalks.com/talks/lessons-from-studying-every-memory-system): Shlok Khemani, Independent. Shlok Khemani examines how consumer AI memory has changed by reverse engineering ChatGPT, Claude, Gemini, and other products. ChatGPT began with a visible list of... - [LLM Knowledge Bases: A Practical Guide](https://aietalks.com/talks/llm-knowledge-bases-a-practical-guide): Ben Holmes, Warp. Ben Holmes presents a workflow for turning scattered personal notes into a browsable knowledge base. He starts with voice dictation because it is faster than... - [Memory Harnesses for Long-Running Research Agents](https://aietalks.com/talks/memory-harnesses-for-long-running-research-agents): Stefania Druga, Sakana.ai. Stefania Druga describes memory as a write, manage, read control loop around a model. Her harness uses small research agents with no durable memory, a... - [Scaling Compute on Context](https://aietalks.com/talks/scaling-compute-on-context): Jack Morris, Engram. Jack Morris frames private knowledge acquisition as a scaling problem. Public models improve by increasing data, training compute, and model size, but those axes mostly... - [Scaling up Continual Learning](https://aietalks.com/talks/scaling-up-continual-learning): Ronak Malde, Trajectory. Ronak Malde argues that continual learning should use the traces produced by models in real applications instead of relying mainly on expensive benchmarks. He compares... - [Agents, codebases, and teams](https://aietalks.com/talks/agents-codebases-and-teams): Aditya Khandelwal, Amazon AGI Lab. Aditya Khandelwal explains why advice for making one repository work well with coding agents often fails when a whole team adopts it. Individual adoption creates... - [The Evolution of Agentic Surfaces](https://aietalks.com/talks/the-evolution-of-agentic-surfaces): Gagan Bhat & Isabella Kai He, Anthropic. Gagan Bhat and Isabella Kai He trace the progression from Anthropic's Messages API to hand-built agent loops, the Claude Agent SDK, and managed agents. As... - [Codex, Behind the Harness](https://aietalks.com/talks/codex-behind-the-harness): Dominik Kundel, OpenAI. Dominik Kundel explains what happens inside the Codex harness after a user sends a message. The app server connects interfaces to the harness, while the... - [Taking Reinforcement Learning Cross Datacenter](https://aietalks.com/talks/taking-reinforcement-learning-cross-datacenter): Nan Jiang, Modal. Nan Jiang argues that reinforcement learning post-training puts two different workloads into the same cluster even though they have different communication needs. Training needs fast... - [Always-on agents run production without the on-call tax](https://aietalks.com/talks/always-on-agents-run-production-without-the-on-call-tax): Justin Smith, Resolve AI. Justin Smith argues that coding agents have increased the flow of changes into production without removing the operational work around those changes. Engineers still spend... - [Guide, Verify, Solve](https://aietalks.com/talks/guide-verify-solve): Anirban Chatterjee, Sonar. Anirban Chatterjee argues that AI-generated code needs a verification process built around zero trust and multiple review methods. A Carnegie Mellon study of GitHub projects... - [Multiplayer agentic engineering](https://aietalks.com/talks/multiplayer-agentic-engineering): Arjun Singh, Superconductor. Arjun Singh describes how Superconductor integrates coding agents into the team's existing human workflows. A single agent session can move between Slack, the app, GitHub,... - [Velocity Sickness: What Happens When Your Whole Team Gets 10x Faster](https://aietalks.com/talks/velocity-sickness-what-happens-when-your-whole-team-gets-10x-faster): Matt Dailey, Ref. Matt Dailey describes velocity sickness as the stress caused by a sudden AI-driven increase in output that produces work without enough impact. He connects it... - [Anthropic's CCA Exam as a Field-Guide for Agentic Engineering](https://aietalks.com/talks/anthropics-cca-exam-as-a-field-guide-for-agentic-engineering): Frank Coyle, UC Berkeley. Frank Coyle uses Anthropic's Claude Certified Architect exam as a practical guide to building agentic systems. He begins with the exam's production scenarios and argues... - [Benchmarking Coding Agents on New vs Legacy Codebases](https://aietalks.com/talks/benchmarking-coding-agents-on-new-vs-legacy-codebases): Denys Linkov, Wisedocs. Denys Linkov describes why Wisedocs refactored an ML pipeline from more than ten legacy repositories into a monorepo. The pipeline processes medical-claims PDFs larger than... - [Realtime multiplayer, automation, and you!](https://aietalks.com/talks/realtime-multiplayer-automation-and-you): Idan Gazit, GitHub. Idan Gazit presents two GitHub Next prototypes built around a shift from individual productivity to group productivity. Agentic Workflows turns a short natural-language request into... - [Compression at the Edge](https://aietalks.com/talks/compression-at-the-edge): Chris Alexiuk, NVIDIA & Daniel Han, Unsloth & Asma Beevi, NVIDIA & Merve Noyan, Hugging Face & Parth Sareen, Ollama. The panel explains why quantization has become necessary as open models grow beyond the memory of ordinary computers. Daniel Han describes shrinking GLM 5.2 from... - [Local Models: Trust, Control, Optimization](https://aietalks.com/talks/local-models-trust-control-optimization): Carter Abdallah, NVIDIA & Vincent Weisser, Prime Intellect & Lucas Atkins, Arcee AI & Chris Alexiuk, NVIDIA. This panel argues that local and open models give builders control over more than model weights. Lucas Atkins defines trust as knowing what you are... - [Open Source Is Dead. Long Live Open Source.](https://aietalks.com/talks/open-source-is-dead-long-live-open-source): Saoud Rizwan, Cline. Saoud Rizwan argues that open source has split into two parts. The contributor community is shrinking because AI-generated pull requests, issues, bug reports, and security... - [The New Primitives: Building AI Native Software](https://aietalks.com/talks/the-new-primitives-building-ai-native-software): Kwindla Kramer, Daily. Kwindla Kramer places current agent development in a longer history of computing. Vannevar Bush's 1945 essay, As We May Think, anticipated technologies including OCR, speech... - [The State of Model Routing](https://aietalks.com/talks/the-state-of-model-routing): Nader Khalil, NVIDIA & Walden Yan, Cognition & Alex Atallah, OpenRouter & Tanay Varshney & Carter Abdallah, NVIDIA. The panel argues that model routing is still an early field, especially for agents whose work changes over time. Walden Yan describes Cognition's Fusion approach:... - [Gadgets: Personal app vibe coding that is actually safe](https://aietalks.com/talks/gadgets-personal-app-vibe-coding-that-is-actually-safe): Kenton Varda, Cloudflare. Kenton Varda argues that personal AI code generation does not fit the standard cloud model. In that model, one developer-owned server runs one approved version... - [Building Turbopuffer](https://aietalks.com/talks/building-turbopuffer): Gergely Orosz, The Pragmatic Engineer & Simon Eskildsen, Turbopuffer. Simon Eskildsen describes how a childhood interest in web development led to competitive programming, Shopify, and eventually Turbopuffer. At Shopify, he worked on sharding, multi-datacenter... - [MCP Apps: Extending the Frontier](https://aietalks.com/talks/mcp-apps-extending-the-frontier): Ido Salomon, MCP steering committee & Liad Yosef, Aura. Ido Salomon and Liad Yosef explain MCP Apps, an extension to MCP for returning interactive interfaces inside chat and coding assistants. A tool call can... - [MCP Tasks (async): Why Aren't Any Agents Supporting Them?](https://aietalks.com/talks/mcp-tasks-async-why-arent-any-agents-supporting-them): Cornelia Davis, Temporal. Cornelia Davis explains why asynchronous MCP tools are harder to implement than their simple interface suggests. A tool call returns a task handle instead of... - [When Will the Benchmaxxing Plague End?](https://aietalks.com/talks/when-will-the-benchmaxxing-plague-end): Nick Heiner, Surge AI. Nick Heiner argues that the word "benchmaxxing" exists because benchmark scores often stop tracking what people want from models. Popular benchmarks can spread through marketing... - [Rethinking Environments for Long-Horizon Work](https://aietalks.com/talks/rethinking-environments-for-long-horizon-work): Rayan Garg, Theta Software. Rayan Garg argues that long-horizon progress depends on how tasks, environments, and verifiers are designed. A task can look long because unrelated work has been... - [Teaching AI to Find Real Vulnerabilities](https://aietalks.com/talks/teaching-ai-to-find-real-vulnerabilities): David Brumley, Bugcrowd. David Brumley explains how to build reinforcement learning environments for cybersecurity. He starts with the way people learn to hack: practice on tasks that become... - [Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It](https://aietalks.com/talks/agents-at-scale-inside-minimaxs-model-and-the-infrastructure-behind-it): Dan Fu, Together AI & Olive Song, MiniMax. Olive Song describes MiniMax's reasons for releasing open-weight models and the post-training work behind MiniMax M3. The model is multimodal, supports applications such as computer... - [Benchmarks: The Good, the Bad, and the Ugly](https://aietalks.com/talks/benchmarks-the-good-the-bad-and-the-ugly): Ali Khial, G2i. Ali Khial explains how coding benchmarks work and why their results can mislead engineers. A benchmark starts with an instruction, sends it to models or... - [Data and Environment Curation for Post-Training LLMs](https://aietalks.com/talks/data-and-environment-curation-for-post-training-llms): Mahesh Sathiamoorthy, Bespoke Labs. Mahesh Sathiamoorthy argues that post-training depends more on the quality and design of its data than on the training algorithm alone. As agents work on... - [Data Quality Is the Compute Multiplier](https://aietalks.com/talks/data-quality-is-the-compute-multiplier): Ari Morcos, DatologyAI. Ari Morcos argues that data quality is the most underused way to get more from limited compute. Better data increases the information each token gives... - [Ending AI Slop](https://aietalks.com/talks/ending-ai-slop): Thais Castello Branco, Taste Labs. Thais Castello Branco argues that AI is much better at coding and math than at design, writing, personality, and emotional intelligence because those domains have... - [Fighting Slop with Slop](https://aietalks.com/talks/fighting-slop-with-slop): Vaibhav Gupta, Boundary. Vaibhav Gupta describes how Boundary builds software without code reviews, fixed AI tooling standards, or a single required workflow. The team writes small architecture rules... - [Learning on the Job: The Future of Post-Training](https://aietalks.com/talks/learning-on-the-job-the-future-of-post-training): Raymond Feng, Applied Compute. Raymond Feng describes post-training systems that let models adapt to the way an enterprise already runs agents. The basic loop has an orchestrator generate rollouts,... - [Reinforcement Learning without Verifiable Rewards](https://aietalks.com/talks/reinforcement-learning-without-verifiable-rewards): Will Brown, Prime Intellect. Will Brown explains how reinforcement learning can extend beyond tasks with exact answers. He starts with a model and harness acting inside an environment, receiving... - [Scaling to Long Horizons](https://aietalks.com/talks/scaling-to-long-horizons): Ross Taylor & Chengxi Taylor, General Reasoning. Ross Taylor opens with Galactica, arguing that curated data, repeated training, and internal thinking tokens anticipated later work on reasoning. His account of early Llama... - [The Base Model Is Dead](https://aietalks.com/talks/the-base-model-is-dead): Varun Singh, Arcee AI. Varun Singh argues that the old base-model recipe has changed. Earlier models were trained mostly on web text, books, and Wikipedia, with post-training used to... - [The Data for Fully Autonomous Software Engineers and Companies](https://aietalks.com/talks/the-data-for-fully-autonomous-software-engineers-and-companies): Joseph Wang, Emulated. Joseph Wang argues that agents will only become reliable autonomous software engineers if their training data includes the difficult work around a codebase. Existing coding... - [Verifiable Environments for AI in Biology](https://aietalks.com/talks/verifiable-environments-for-ai-in-biology): Kenny Workman, LatchBio. Kenny Workman describes LatchBio's effort to build agents and benchmarks for scientific biology workflows. Modern single-cell and spatial experiments can produce terabytes of data, but... - [What's Next After RLHF?](https://aietalks.com/talks/whats-next-after-rlhf): Diogo Almeida, TypeSafe AI. Diogo Almeida argues that the field is confusing assistance with automation. RLHF trains models to optimize for human preferences, so they are good at producing... - [Build for the Memo, Not the Demo](https://aietalks.com/talks/build-for-the-memo-not-the-demo): Shawn Chan, China Resources Holdings. Shawn Chan compares an impressive AI demo with an investment memo. A demo produces a fluent answer that makes people interested for a few minutes.... - [First Steps Toward Automated AI Research](https://aietalks.com/talks/first-steps-toward-automated-ai-research): Richard Socher, Recursive AI. Richard Socher presents automated scientific research as his version of going to Mars. He draws on biological evolution and Karl Popper's account of science as... - [Let's integrate AI Agents in Event-Sourced Systems](https://aietalks.com/talks/lets-integrate-ai-agents-in-event-sourced-systems): Divakar Kumar, FlyersSoft. Divakar Kumar describes an agentic layer for real-time fraud detection that sits alongside an existing rule-based engine and ML model. Those systems approve or reject... - [Your Finance Agent's Bottleneck Is You](https://aietalks.com/talks/your-finance-agents-bottleneck-is-you): Ramana Siddanth Emani, Auditoria AI. Ramana Siddanth Emani argues that production finance agents fail less because of model capability than because developers cannot move through the bug-fixing loop fast enough.... - [How Kepler Built Verifiable AI for Financial Services](https://aietalks.com/talks/how-kepler-built-verifiable-ai-for-financial-services): Vinoo Ganesh, Kepler. Vinoo Ganesh argues that financial AI has a verification problem. Language models are useful for reading, reasoning, and planning, but they are probability machines and... - [Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains](https://aietalks.com/talks/morgan-stanleys-alphalab-multi-agent-research-across-optimization-domains): Brendan Rappazzo, Morgan Stanley. Brendan Rappazzo presents AlphaLab, Morgan Stanley's open-source system for automating quantitative research. A user supplies data access and a prediction goal in natural language. AlphaLab... - [Persona Engineering: A Field Guide to AI Synthetic Personas](https://aietalks.com/talks/persona-engineering-a-field-guide-to-ai-synthetic-personas): Ishan Anand, InsightSciences.ai. Ishan Anand presents synthetic personas as forecasts of people rather than substitutes for people. He describes a study in which agents reproduced survey and personality-test... - [SimulationMaxxing: How We Ship Agents 20× Faster](https://aietalks.com/talks/simulationmaxxing-how-we-ship-agents-20-faster): Aman Gupta, Nubank & Shreya Rajpal, Snowglobe. Aman Gupta and Shreya Rajpal describe how Nubank uses Snowglobe simulations to evaluate customer-support agents before exposing changes to live users. Agent evaluation data is... - [Skills are New Features: Building a Skill-Centric Harness](https://aietalks.com/talks/skills-are-new-features-building-a-skill-centric-harness): Yogendra Miraje, FactSet. Yogendra Miraje describes how FactSet moved from its own blueprint format to open-source skills and began treating skills as the feature layer of agentic products.... - [We Vetted 2,000 AI Skills Before They Reached Developers](https://aietalks.com/talks/we-vetted-2-000-ai-skills-before-they-reached-developers): Lucas Palma, Nubank. Lucas Palma describes how Nubank reviews AI skills before engineers can get them from its internal marketplace. He treats a skill as a supply-chain dependency... - [Wearing the Agent: From Group Chats to Glasses](https://aietalks.com/talks/wearing-the-agent-from-group-chats-to-glasses): Sai Krishna Rallabandi. Sai Krishna Rallabandi describes eight months of running Judith in a group of friends and family, then connects that experience to agents in group chats... - [Why Off-the-Shelf AI Doesn't Understand Money](https://aietalks.com/talks/why-off-the-shelf-ai-doesnt-understand-money): Udi Menkes, Intuit. Udi Menkes argues that general-purpose language models often bluff when answering financial questions. They know what people have written about money, but they have not... - [Your Agent Didn't Fail. Your Harness Did.](https://aietalks.com/talks/your-agent-didnt-fail-your-harness-did): Vinoth Govindarajan, OpenAI. Vinoth Govindarajan argues that many production agent failures are system failures rather than model failures. An agent can send a successful reply while failing to... - [AI Agents for Performance: Ship Faster, Pay Less](https://aietalks.com/talks/ai-agents-for-performance-ship-faster-pay-less): Rajat Shah, Netflix. Rajat Shah describes a workflow that uses an AI coding agent to turn production profiling data into a proposed performance fix. The agent reads a... - [AI tools for Forward Deployed Engineering](https://aietalks.com/talks/ai-tools-for-forward-deployed-engineering): Vasuman Moza & JD Pruitt, Varick Agents. Vasuman Moza argues that AI has moved past the execution bottleneck. Models can now perform many knowledge-work tasks, but they still struggle to understand how... - [Forward Deployed Engineering 101](https://aietalks.com/talks/forward-deployed-engineering-101): Kevin Bai, Anthropic. Kevin Bai explains forward deployed engineering through Palantir's use of Foundry. Foundry centralizes an organization's data, creates a shared ontology, and lets customers build applications,... - [How Forward Deployed Engineering is Done at Cognition](https://aietalks.com/talks/how-forward-deployed-engineering-is-done-at-cognition): Jia Wu, Cognition. Jia Wu explains how Cognition deploys Devin inside large enterprise customers. The FDE team starts by understanding the customer's software work, including backlogs, migrations, testing,... - [How Forward Deployed Engineering Is Done at Decagon](https://aietalks.com/talks/how-forward-deployed-engineering-is-done-at-decagon): Sunny Rekhi, Decagon. Sunny Rekhi explains how forward deployed engineering works at Decagon, which provides an AI customer service agent for calls, email, and other channels. The work... - [How Forward Deployed Engineering is done at Factory](https://aietalks.com/talks/how-forward-deployed-engineering-is-done-at-factory): Eno Reyes, Factory. Eno Reyes describes forward deployed engineering at Factory as product engineering carried out close to major customers. The deployed engineer learns how an organization builds... - [How Forward Deployed Engineering is done at Kepler](https://aietalks.com/talks/how-forward-deployed-engineering-is-done-at-kepler): Vinoo Ganesh, Kepler. Vinoo Ganesh argues that forward deployed engineering should belong to product strategy rather than sales or customer success. The engineer works inside the customer's environment,... - [How Forward Deployed Engineering is done at Ramp](https://aietalks.com/talks/how-forward-deployed-engineering-is-done-at-ramp): Leo Mehr, Ramp. Leo Mehr describes two principles from building Forward Deployed Engineering at Ramp. The first is to always be scoping. FDEs should understand what is driving... - [Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub](https://aietalks.com/talks/serving-2-million-models-without-melting-scaling-the-hugging-face-hub): Arek Borucki, Hugging Face. Arek Borucki explains how Hugging Face operates the Hub as its catalog grew to 3 million public models, 1 million datasets, and more than 14... - [The Dirty Secret of Forward Deployed Engineering](https://aietalks.com/talks/the-dirty-secret-of-forward-deployed-engineering): Natalie Meurer, Sierra. Natalie Meurer traces forward deployed engineering from its origins at Palantir. In the early years, the job meant sitting with customers, often at on-premise deployments,... - [DeepSWE: A Contamination-Resistant Coding Benchmark](https://aietalks.com/talks/deepswe-a-contamination-resistant-coding-benchmark): James Shi, Datacurve. James Shi introduces DeepSWE, Datacurve's long-horizon software engineering benchmark. It contains 113 original tasks across 91 active open-source repositories, written by engineers who understand those... - [State of Data](https://aietalks.com/talks/state-of-data): Sean Cai, Independent. Sean Cai argues that the valuable part of the data market is moving beyond image labels and other static records. Models need process data from... - [The Messy Reality of Scale: Synthetic Data and Pre-Training](https://aietalks.com/talks/the-messy-reality-of-scale-synthetic-data-and-pre-training): Marah Abdin & Robert McHardy, poolside. Marah Abdin describes synthetic data as a way to complement organic data. Poolside uses it to expose implicit rationale, planning, and structure, fill gaps, reduce... - [Evaling Video Slop](https://aietalks.com/talks/evaling-video-slop): Maor Bril, Character.ai. Maor Bril explains why evaluating generated video needs different methods from evaluating text or still images. CLIP can check whether a frame matches a prompt,... - [Evals-Driven Development for a Mental Health AI Coach](https://aietalks.com/talks/evals-driven-development-for-a-mental-health-ai-coach): Akele Reed, Dave Revere & Doug Keller, SonderMind. Akele Reed and Dave Revere describe how SonderMind built a mental health AI coach around clinical review and repeated evaluation. The system uses separate input... - [From Agent Traces to Agent Simulations](https://aietalks.com/talks/from-agent-traces-to-agent-simulations): Rustem Feyzkhanov, Snorkel AI. Rustem Feyzkhanov argues that teams need private benchmarks built from their own production traces. A trace becomes a replayable task when the team reconstructs the... - [Loop Engineering from First Principles](https://aietalks.com/talks/loop-engineering-from-first-principles): Kyle Mistele, HumanLayer. Kyle Mistele argues that coding agents need better-engineered loops, not simply better prompts. A blind loop can generate a 40,000-line pull request that nobody wants... - [Why Large? Tiny LMs and Agents on Edge and Robotics](https://aietalks.com/talks/why-large-tiny-lms-and-agents-on-edge-and-robotics): Cormac Brick, Google. Cormac Brick explains why edge AI needs models that fit the device's RAM budget. Cloud inference brings latency, connectivity, privacy, and cost problems, while DRAM... - [Building Closed-Loop Evals for a Multimodal Agent at Scale](https://aietalks.com/talks/building-closed-loop-evals-for-a-multimodal-agent-at-scale): Soumya Gupta & Jai Chopra, Uber. Soumya Gupta and Jai Chopra describe Uber's production system for improving food photography from smaller independent Uber Eats merchants. The system first understands and routes... - [Everything Is a Rollout](https://aietalks.com/talks/everything-is-a-rollout): Alex Shaw & Ryan Marten, Laude Institute. Alex Shaw presents agent development as a machine learning workflow. Traditional software engineering often lets a developer predict what code will do before running it,... - [From Signal to PR: Anatomy of a Self-Improving Agent](https://aietalks.com/talks/from-signal-to-pr-anatomy-of-a-self-improving-agent): Jason Lopatecki, Arize. Jason Lopatecki describes Arize's move from Alyx, its earlier product agent, to Signal, a system that investigates production problems before a human starts debugging. The... - [How Evals and Prompts Shape Agent Behavior](https://aietalks.com/talks/how-evals-and-prompts-shape-agent-behavior): Preetika Bhateja, Google/YouTube & Daniel Bump & Chris Souza, Google. Preetika Bhateja and Daniel Bump describe how the YouTube Ads team built an evaluation workflow for an agent that turns messy advertising creatives into reusable... - [Setting Yourself Up for Success](https://aietalks.com/talks/setting-yourself-up-for-success): Jason Liu, OpenAI. Jason Liu presents Codex as a general work system rather than only a coding assistant. His setup starts with a personal monorepo containing project and... - [The Future of Evals: From LLM as a Judge to Agent as a Judge](https://aietalks.com/talks/the-future-of-evals-from-llm-as-a-judge-to-agent-as-a-judge): Aparna Dhinakaran, Arize AI. Aparna Dhinakaran argues that evals have to evolve with the systems they measure. Early agents mainly answered prompts, while newer agents use tools, reason over... - [Training Frontier Models to Out-Think Hackers](https://aietalks.com/talks/training-frontier-models-to-out-think-hackers): Uri Rolls, Arithmetic & Thom Wolf, Hugging Face. Uri Rolls and Thom Wolf present Arithmetic's first cybersecurity benchmark, focused on access-control failures. The benchmark puts models inside realistic blackbox environments built from chained... - [Vending-Bench: Long-Horizon Agent Evals](https://aietalks.com/talks/vending-bench-long-horizon-agent-evals): Lukas Petersson, Andon Labs. Lukas Petersson describes Andon Labs' effort to test AI agents over long periods by giving them businesses to run. Vending Bench places models in a... - [AI on Your Lakehouse: Context Comes in Shapes, Not Queries](https://aietalks.com/talks/ai-on-your-lakehouse-context-comes-in-shapes-not-queries): Zach Blumenfeld, Neo4j. Zach Blumenfeld presents a hands-on workshop for giving agents better context over lakehouse data. The example combines BigQuery tables with manuals, bulletins, recalls, and work-order... - [Citation Needed: Provenance for LLM-Built Knowledge Graphs](https://aietalks.com/talks/citation-needed-provenance-for-llm-built-knowledge-graphs): Daniel Chalef, Zep AI. Daniel Chalef argues that provenance must be part of the data model for knowledge graphs built by LLM pipelines. An LLM may merge identities, combine... - [Harness Engineering is Not Enough: Why Software Factories Fail](https://aietalks.com/talks/harness-engineering-is-not-enough-why-software-factories-fail): Dex Horthy, HumanLayer. Dex Horthy argues that software factories fail when they remove humans from code review before models can preserve a codebase's quality. Coding models are trained... - [Learned Execution Graphs for Anomaly Detection & Drift in APIs](https://aietalks.com/talks/learned-execution-graphs-for-anomaly-detection-drift-in-apis): Ritvik Pandya, JP Morgan Chase. Ritvik Pandya presents execution graphs as a way to inspect how an API request moves through gateways, authentication, orchestration, and downstream services. Each request becomes... - [Local Agentic Theory For Mobile Games](https://aietalks.com/talks/local-agentic-theory-for-mobile-games): Shafik Quoraishee & Joanne Song, The New York Times. Shafik Quoraishee and Joanne Song present experimental work on agents that run inside mobile games and accessibility systems. They distinguish reinforcement learning, which changes model... - [Notion's Token Town](https://aietalks.com/talks/notions-token-town): Sarah Sachs, Notion. Sarah Sachs explains how Notion manages the cost and dependency risks of building AI products. Model upgrades can keep the same token price while using... - [Perception Agents](https://aietalks.com/talks/perception-agents): Antje Barth, Amazon AGI Lab. Antje Barth argues that computer-use agents have learned to click, type, scroll, call APIs, and run workflows, but still struggle with end-to-end work such as... - [The Unreasonable Effectiveness of Separating the Task from the Model](https://aietalks.com/talks/the-unreasonable-effectiveness-of-separating-the-task-from-the-model): Maxime Rivest, DSPy & Isaac Miller, DSPy; cmpnd. Maxime Rivest and Isaac Miller argue that AI programs should be built like reusable software functions. A task gets a defined input and output interface,... - [Video Has No Memory. Here's How We Built One.](https://aietalks.com/talks/video-has-no-memory-heres-how-we-built-one): James Le, TwelveLabs. James Le argues that most video AI systems have no memory in the system sense. They sample frames, extract transcripts, or place content in a... - [Why Agentic Systems Need Ontologies](https://aietalks.com/talks/why-agentic-systems-need-ontologies): Frank Coyle, UC Berkeley. Frank Coyle argues that agent failures often come from giving a probabilistic model too much freedom over a domain it only partly understands. Prompts written... - [Why We Killed Our Multi-Agent Pipeline](https://aietalks.com/talks/why-we-killed-our-multi-agent-pipeline): Subbiah Sethuraman & Abhilash Asokan, ZS Associates. Subbiah Sethuraman and Abhilash Asokan describe how ZS Associates replaced a multi-agent pharma analytics pipeline with a smaller architecture. The original system copied an analyst's... - [Active Graph Agent Runtime (BabyAGI 4)](https://aietalks.com/talks/active-graph-agent-runtime-babyagi-4): Yohei Nakajima, Untapped Capital. Yohei Nakajima presents ActiveGraph, an experimental open-source runtime for building agents around an immutable event log rather than around an LLM session. Every agent action... - [Claude for Long-Horizon Tasks](https://aietalks.com/talks/claude-for-long-horizon-tasks): Lance Martin, Anthropic. Lance Martin describes how Anthropic is building agents that can work for much longer without constant human steering. He traces the shift from short model... - [CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens](https://aietalks.com/talks/crabrag-why-automated-assistants-need-graph-memory-not-more-tokens): Stephen Chin, Neo4j. Stephen Chin argues that agent memory needs explicit relationships, especially when an assistant works with a large and connected body of information. He shows how... - [From Systems of Record to Systems of Context](https://aietalks.com/talks/from-systems-of-record-to-systems-of-context): Omri Bruchim & Tomer Ast, monday.com. Omri Bruchim and Tomer Ast argue that the main problem with work assistants is not retrieval. An assistant may have access to boards, tasks, emails,... - [Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer](https://aietalks.com/talks/thinner-agents-on-a-smarter-substrate-the-ontology-based-semantic-layer): Emil Eifrem, Neo4j. Emil Eifrem describes a common failure pattern in enterprise agent development. A bank account-opening agent may need DMV records and passport verification, but every new... - [Your Moat Is Your Data Model](https://aietalks.com/talks/your-moat-is-your-data-model): Mike Phipps, Gates Foundation. Mike Phipps argues that models, chat interfaces, and agent frameworks are becoming easier to replace, while an organization's understanding of its own processes remains harder... - [Better Auth](https://aietalks.com/talks/better-auth): Bereket Habtemeskel & Paola Estefania, Better Auth. Paola Estefania presents Agent Auth, a protocol co-created with Bereket Habtemeskel for agents that act for a user or organization. She starts with the risk... - [Every Harness Will Become A Claw](https://aietalks.com/talks/every-harness-will-become-a-claw): Sam Bhagwat, Mastra. Sam Bhagwat describes a progression from LLMs to agents, harnesses, and claws. An agent has an action loop, tool calls, memory, retries, and context engineering.... - [HTML Is All Agents Need](https://aietalks.com/talks/html-is-all-agents-need): James Russo, HeyGen. James Russo explains how HeyGen built Hyperframes, an open-source framework that turns HTML into video. The team tried tools including Remotion, but found that teaching... - [The Biggest Challenge in Your Stack? Evals, Evals, Evals](https://aietalks.com/talks/the-biggest-challenge-in-your-stack-evals-evals-evals): Barr Yaron, Amplify Partners. Barr Yaron presents results from a survey of 1,048 people involved in AI engineering. She describes a field that crosses job titles and company sizes,... - [The Desktop Frontier](https://aietalks.com/talks/the-desktop-frontier): Ahmad Osman, Osmantic. Ahmad Osman argues that frontier-level AI is moving onto personal hardware because model efficiency is improving faster than model size is growing. He uses the... - [Your agent architecture has a half-life of 6 months](https://aietalks.com/talks/your-agent-architecture-has-a-half-life-of-6-months): Dan Farrelly, Inngest. Dan Farrelly argues that agent architectures decay because teams couple fast-changing parts of the system to slower, more stable parts. He divides an agent into... - [Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing](https://aietalks.com/talks/agent-output-is-not-ux-rendering-layer-your-llm-pipeline-is-missing): Bala Ramdoss, Amazon. Bala Ramdoss argues that model output is only the raw material for an AI product experience. A useful feature needs a rendering layer between the... - [Build the AI GTM Agent That Knows the Buyer](https://aietalks.com/talks/build-the-ai-gtm-agent-that-knows-the-buyer): Dr. Sajjan Kanukolanu, Position2 (Position Squared). Dr. Sajjan Kanukolanu argues that most AI GTM deployments start at the wrong point. They add a chatbot to an existing stack, then ask visitors... - [Can Oncology Workflows Run Without Human Touch?](https://aietalks.com/talks/can-oncology-workflows-run-without-human-touch): Anant Shankhdhar, Risa Labs. Anant Shankhdhar describes Risa Labs' attempt to remove human review from parts of oncology prior authorization. The workflow starts by collecting patient and insurance information... - [Designing Voice Agents for Real Conversations](https://aietalks.com/talks/designing-voice-agents-for-real-conversations): Chintan Agrawal & Daniel Wirjo, AWS. Chintan Agrawal and Daniel Wirjo explain why voice agents need different engineering from chat agents. Human turn switches happen in about 200 milliseconds, while delays... - [Don't Let the LLM Drive](https://aietalks.com/talks/dont-let-the-llm-drive): Ornella Bahidika, Microsoft & Joel Allou. Ornella Bahidika and Joel Allou describe Ace, a live voice tutor that runs lessons through a harness rather than letting the language model control the... - [Enterprise Agents Have a Structure Problem](https://aietalks.com/talks/enterprise-agents-have-a-structure-problem): Ishita Daga, Tesla. Ishita Daga argues that enterprise agents often fail because they lack structure around business data. Bigger models, longer prompts, Markdown files, and more tools do... - [In the Land of AI Agents, the Verifiers Are King](https://aietalks.com/talks/in-the-land-of-ai-agents-the-verifiers-are-king): Tariq Shaukat, Sonar. Tariq Shaukat argues that the hard part of enterprise AI is no longer generating plausible output. It is determining whether that output is correct and... - [It's 10pm. Do You Know Where Your Agents Are?](https://aietalks.com/talks/its-10pm-do-you-know-where-your-agents-are): Kim Maida, Keycard. Kim Maida shows how an incident-management agent can misuse a kitchen-sink API key while handling routine night-shift tickets. In the demo, the agent drops a... - [Medic for Apache Spark: First Aid for Failing Jobs](https://aietalks.com/talks/medic-for-apache-spark-first-aid-for-failing-jobs): Drasko Profirovic, Pinterest. Drasko Profirovic describes Medic, an agentic tool that investigates failed Apache Spark jobs and produces evidence-based diagnoses with suggested fixes. The first prototype connected an... - [Security Track Intro](https://aietalks.com/talks/security-track-intro): Randall Degges, Snyk. Randall Degges opens the World's Fair Security Track by describing the security problems that limit AI-assisted software development. AI helps developers ship better software faster,... - [Skills are the New SDKs](https://aietalks.com/talks/skills-are-the-new-sdks): Elvin Aghammadzada, DataRobot. Elvin Aghammadzada argues that platforms need a skill layer if coding agents are going to use them reliably. APIs, SDKs, MCP tools, and human-focused documentation... - [We Gave an Agent Production Code Access and Then Tried to Sleep at Night](https://aietalks.com/talks/we-gave-an-agent-production-code-access-and-then-tried-to-sleep-at-night): Moritz Johner, Form3. Moritz Johner describes PatchPilot, a production system Form3 built to remediate CVEs across thousands of repositories. Standard dependency tools can update manifests, but they miss... - [When Agents Meet Physical Data: The Other Physics of Agent Harnesses](https://aietalks.com/talks/when-agents-meet-physical-data-the-other-physics-of-agent-harnesses): Dmitry Petrov, DataChain. Dmitry Petrov explains why coding-agent habits fail on large collections of video, sensor data, and robot telemetry. Raw files contain nested structures such as clips,... - [Why Your Agent Disagrees With Itself (And What To Do About It)](https://aietalks.com/talks/why-your-agent-disagrees-with-itself-and-what-to-do-about-it): Diane Lin, Datadog. Diane Lin explains why the same AI agent can give materially different answers for the same input. In sentiment analysis and cybersecurity alert triage, these... - [Your LLM Stack Is a 2008 Database With Better Marketing](https://aietalks.com/talks/your-llm-stack-is-a-2008-database-with-better-marketing): Lovina Dmello, NVIDIA. Lovina Dmello argues that production ML security has a familiar source: infrastructure mistakes. She opens with thousands of Ray clusters exposed on the public internet... - [Your Voice Agent Doesn't Need a Frontier Model](https://aietalks.com/talks/your-voice-agent-doesnt-need-a-frontier-model): Joel Allou & Ornella Bahidika, Microsoft. Joel Allou and Ornella Bahidika describe Ace, a live AI voice tutor built around a small model. Their model choice comes from the timing of... - [Build Evals That Actually Matter](https://aietalks.com/talks/build-evals-that-actually-matter): Nick Ung & Akshay Sharma, Lyft. Nick Ung and Akshay Sharma describe Lyft's evaluation pipeline for customer-support agents. They separate development, offline testing, production tracing, online grading, and human error analysis.... - [From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization](https://aietalks.com/talks/from-blind-spots-to-merged-prs-continuous-agentic-performance-optimization): May Walter, Hud. May Walter describes a workflow that uses coding agents to find and verify performance improvements in live production systems. The problem is the unpredictable investigation... - [From Tokens to Cells: Foundation Models for Single-Cell Biology](https://aietalks.com/talks/from-tokens-to-cells-foundation-models-for-single-cell-biology): Akram Baharlouei, Altos Labs. Akram Baharlouei explains foundation models for single-cell biology from the perspective of a machine learning engineer without a biology background. He starts with cellular reprogramming... - [You Didn't Ship a Bug. You Just Wrote It for a Human.](https://aietalks.com/talks/you-didnt-ship-a-bug-you-just-wrote-it-for-a-human): Ravi Madabhushi, Scalekit. Ravi Madabhushi describes a production incident caused by a human-centered assumption. Scalekit's system updated a user's last-seen timestamp on every tool call. An agent called... - [A Practitioner's Guide to Graphs](https://aietalks.com/talks/a-practitioners-guide-to-graphs): Tim Ainge, Good Collective. Tim Ainge gives a practical introduction to graphs for people building AI applications. He starts with nodes, edges, labels, properties, and direction, then focuses on... - [Agents Need Feature Flags](https://aietalks.com/talks/agents-need-feature-flags): Sachin Gupta. Sachin Gupta argues that agent systems are being deployed with far less operational control than ordinary web software. A prompt edit, model swap, new tool,... - [Agents Need Receipts, Not More Tool Calls](https://aietalks.com/talks/agents-need-receipts-not-more-tool-calls): Armanas Povilionis, Alithea Bio. Armanas Povilionis argues that agents need a reliable way to coordinate work across organizational boundaries. More tools improve what an agent can do inside one... - [Autonomous Agents for Scientific Tasks](https://aietalks.com/talks/autonomous-agents-for-scientific-tasks): Sina Shahandeh, Radicait. Sina Shahandeh argues that autonomous agents for scientific work need more than the hill-climbing loops used in coding benchmarks and toy ML tasks. Those agents... - [Content Is Code](https://aietalks.com/talks/content-is-code): Matt Palmer, Conductor. Matt Palmer argues that code is becoming the main medium for producing technical content. This includes documentation, change logs, product updates, websites, videos, slides, motion... - [Stop Burning Tokens: Why Self-Improvement Needs Domain Expertise First](https://aietalks.com/talks/stop-burning-tokens-why-self-improvement-needs-domain-expertise-first): Annabell Schäfer, Langfuse. Annabell Schäfer argues that self-improvement loops are only as useful as the target function they optimize. Coding provides an unusually clear signal because code either... - [Stop Renting Your Cognitive Infrastructure](https://aietalks.com/talks/stop-renting-your-cognitive-infrastructure): Thiyagarajan Maruthavanan, Kalmantic Labs. Thiyagarajan Maruthavanan argues that teams lose control when they treat inference as a cheap, metered service. He describes his own costs rising after an app... - [The UX of AI: Making AI-Powered Apps Your Users Don't Hate](https://aietalks.com/talks/the-ux-of-ai-making-ai-powered-apps-your-users-dont-hate): Kathryn Grayson Nanz, Progress Software. Kathryn Grayson Nanz argues that AI-powered software has a user experience problem because developers understand concepts that most users have not learned yet. She compares... - [Your Agents Need a Save Button](https://aietalks.com/talks/your-agents-need-a-save-button): Hamza Tahir, ZenML. Hamza Tahir argues that agent infrastructure needs a durable checkpoint connected to the runtime, rather than relying on read-only traces. A checkpoint captures state around... - [Every company should have a Brain](https://aietalks.com/talks/every-company-should-have-a-brain): Garry Tan, Y Combinator. Garry Tan argues that AI-native companies can let small teams perform work that previously required much larger organizations. He bases this on his own shift... - [On AI and Knowledge](https://aietalks.com/talks/on-ai-and-knowledge): Pablo Castro, Microsoft. Pablo Castro presents AI applications through three kinds of knowledge. Intrinsic knowledge is the information stored in a model's parametric memory, which powered early tools... - [Software engineering is not about writing code](https://aietalks.com/talks/software-engineering-is-not-about-writing-code): Benoit Schillings, Google DeepMind. Benoit Schillings argues that software development has entered an AI frontier where generating ordinary code is no longer the main difficulty. Earlier eras were limited... - [Special Topics in Kernels, RL, Reward Hacking in Agents](https://aietalks.com/talks/special-topics-in-kernels-rl-reward-hacking-in-agents): Daniel Han, Unsloth. Daniel Han gives a broad technical seminar on current model progress, open and closed models, inference, benchmarks, kernels, reinforcement learning, and agent behavior. He argues... - [The Great Loops Debate](https://aietalks.com/talks/the-great-loops-debate): Ali Howard & Ian Livingstone, Keycard & Geoff Huntley & Greg Pstrucha, Sentry & Dex Horthy, HumanLayer. The debate asks whether the excitement around agent loops matches what works in practice. Geoff Huntley and Ian Livingstone argue that loops are already useful... - [Using LLMs to Secure Source Code](https://aietalks.com/talks/using-llms-to-secure-source-code): Eugene Yan, Anthropic. Eugene Yan describes how Anthropic has worked with security teams to find and fix vulnerabilities in source code. He argues that models have improved enough... - [Imagination Engineering: "Live in the future and then build what's missing."](https://aietalks.com/talks/imagination-engineering-live-in-the-future-and-then-build-whats-missing): Eve Bouffard, Y Combinator. Eve Bouffard argues that increasingly capable AI models will make execution much easier, so people should spend more effort stretching their imagination and inventing things... - [Computer-Use 2.0: Agents Just Got Multi-Cursor](https://aietalks.com/talks/computer-use-2-0-agents-just-got-multi-cursor): Francesco Bonacci, Cua. Francesco Bonacci and his team describe a move beyond the usual computer-use loop, where an agent repeatedly reads a full screenshot and controls the foreground... - [Recursive Model Improvement](https://aietalks.com/talks/recursive-model-improvement): Lee Robinson, Cursor. Lee Robinson explains how Cursor trains its models through two connected loops. The outer loop collects product feedback, internal reports, online metrics, and A/B test... - [Don't Ship Skills Without Evals](https://aietalks.com/talks/dont-ship-skills-without-evals): Philipp Schmid, Google DeepMind. Philipp Schmid explains why agent skills need tests before they reach users. A skill usually contains a short description, a fuller instruction file, and deeper... - [Forward Deployed Engineering at Cursor](https://aietalks.com/talks/forward-deployed-engineering-at-cursor): Pauline Brunet, Cursor. Pauline Brunet describes how Cursor decides when forward deployed engineering is useful and how it runs the work. The first test is the customer's maturity... - [The engineer of the future is the person who is able to choose what is worth doing.](https://aietalks.com/talks/the-engineer-of-the-future-is-the-person-who-is-able-to-choose-what-is-worth): Addy Osmani. Addy Osmani argues that AI agents will automate more coding while making engineering judgment more important. Agents can investigate, implement, test, and report inside an... - [WTF Is the Context Layer? The Missing Infrastructure for Production Agents](https://aietalks.com/talks/wtf-is-the-context-layer-the-missing-infrastructure-for-production-agents): Prukalpa Sankar, Atlan. Prukalpa Sankar argues that model intelligence has advanced much faster than the business context agents need to work reliably. A human analyst can answer a... - [From fork() to Fleet: Designing an Agent Sandbox Cloud](https://aietalks.com/talks/from-fork-to-fleet-designing-an-agent-sandbox-cloud): Abhishek Bhardwaj, OpenAI. Abhishek Bhardwaj explains why agent products need secure computers that can execute untrusted model-generated code. He starts with Linux execution and compares fork() and exec,... - [I've Never Seen Anything Scarier Than an LLM with Tool Calls](https://aietalks.com/talks/ive-never-seen-anything-scarier-than-an-llm-with-tool-calls): Erik Meijer, LiveNet Labs. Erik Meijer argues that language models are dangerous because they pursue goals through imperfect instructions, prompt injection, and tool calls that can alter the real... - [Modern Post-Training: A Deep Dive](https://aietalks.com/talks/modern-post-training-a-deep-dive): Will Brown, Prime Intellect. Will Brown presents Prime Intellect's open-source approach to modern post-training. He treats environments as reusable specifications of data, interaction, and scoring. The same environment can... - [Stop Evaluating Models Like It's the 50s](https://aietalks.com/talks/stop-evaluating-models-like-its-the-50s): Alejandro Vidal, Mindmakers. Alejandro Vidal argues that LLM evaluation still relies too heavily on Classical Test Theory: count correct answers, divide by the number of questions, and treat... - [A Song of Types and Agents](https://aietalks.com/talks/a-song-of-types-and-agents): Roberto Stagi, Ratel. Roberto Stagi argues that Python and TypeScript are winning different parts of the AI stack. Python remains the language for training, research, and GPU serving.... - [remobi.app: Don't Change Your Terminal Workflow for Mobile](https://aietalks.com/talks/remobi-app-dont-change-your-terminal-workflow-for-mobile): Connor Adams. Connor Adams presents Remobi, an open-source progressive web app for controlling a remote terminal from a phone without changing the way a developer works. His... - [ReviewDebt: A Practical Framework for Scoring Every Pull Request](https://aietalks.com/talks/reviewdebt-a-practical-framework-for-scoring-every-pull-request): Sachin Gupta, Ebay. Sachin Gupta argues that coding agents are increasing code production faster than humans can review it. The resulting review debt is the accumulated gap between... - [Semantic Blindness: 500,000 Sensors Confused an LLM](https://aietalks.com/talks/semantic-blindness-500-000-sensors-confused-an-llm): Raahul Singh & Vanč Levstik, Phaidra. Raahul Singh and Vanč Levstik describe deploying LLMs against AI-factory and data-center systems with hundreds of thousands of sensors, equipment nodes, and inconsistent naming conventions.... - [The Agentic Web and the Bazaar Era of AI](https://aietalks.com/talks/the-agentic-web-and-the-bazaar-era-of-ai): Ramesh Raskar, MIT Media Lab. Ramesh Raskar presents Project Nanda as an open infrastructure project for a web of AI agents. He compares today's agent ecosystem to the AOL era,... - [The AI bugpocalypse is here. Now what?](https://aietalks.com/talks/the-ai-bugpocalypse-is-here-now-what): Jack Cable, Corridor. Jack Cable argues that defenders face two changes at once: AI models are getting better at finding and exploiting vulnerabilities, while coding agents are producing... - [What Does Done Even Mean? Agents and Paperclip's Liveness Model](https://aietalks.com/talks/what-does-done-even-mean-agents-and-paperclips-liveness-model): Dotta, Paperclip. Dotta argues that an agent changing a task to done does not settle whether the work is ready to merge, deploy, or announce. Completion includes... - [Chat and citations won't save your vertical AI](https://aietalks.com/talks/chat-and-citations-wont-save-your-vertical-ai): Atul Ramachandran, Filed Inc. Atul Ramachandran argues that chat and citations do not remove enough work from users of vertical AI products. Chat keeps the interaction synchronous, while citations... - [Claws Out: Securing and Building with OpenClaw](https://aietalks.com/talks/claws-out-securing-and-building-with-openclaw): Nick Taylor, Pomerium. Nick Taylor explains how he secured his OpenClaw instance with the trusted proxy auth mode he contributed to the project. Before this mode, users still... - [Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers](https://aietalks.com/talks/design-patterns-for-ai-trust-juries-libraries-and-agent-tiers): Alex Bauer, Upside.tech. Alex Bauer argues that the main AI problem for go-to-market teams has become trust. Agents make it easy for people without engineering backgrounds to build... - [Develop at Idea Velocity](https://aietalks.com/talks/develop-at-idea-velocity): Jeffrey Lee-Chan, Snapchat. Jeffrey Lee-Chan describes a development setup built around OpenClaw, parallel agents, tmux terminals, and separate work trees. He sends short messages because OpenClaw retains context... - [Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD](https://aietalks.com/talks/every-solo-agent-builder-eventually-reinvents-a-worse-version-of-ci-cd): Sumaiya Shrabony. Sumaiya Shrabony uses her open-source, 19-skill Claude Code agent system as a case study for the operational controls agent builders tend to recreate badly. The... - [From Writing Code to Designing Systems: How the Developer Role is Changing](https://aietalks.com/talks/from-writing-code-to-designing-systems-how-the-developer-role-is-changing): Chris Noring, Microsoft. Chris Noring argues that AI has changed the developer's center of gravity. Developers still write code, but they spend less time producing every line and... - [State of the Union: Why Local, Why Now](https://aietalks.com/talks/state-of-the-union-why-local-why-now): Nader Khalil, NVIDIA & Joseph Nelson, Roboflow & Alex Cheema, EXO Labs & Ahmad Osman, Osmantic & Matthew Berman, Forward Future. The panel argues that local AI reached an inflection point because capable models now fit on phones, desktops, and small clusters, while harnesses give them... - [Stop AI Agent Hallucinations: 5 Techniques + Production Patterns](https://aietalks.com/talks/stop-ai-agent-hallucinations-5-techniques-production-patterns): Elizabeth Fuentes, AWS. Elizabeth Fuentes presents five code-level techniques for reducing hallucinations and wasted tokens in AI agents. Semantic tool selection filters a travel agent's 29 tools down... - [The Factory That Dreams: 39 AI Agents, No Framework](https://aietalks.com/talks/the-factory-that-dreams-39-ai-agents-no-framework): Rushabh Doshi, Machinecraft / Fork My Brain. Rushabh Doshi describes how Machinecraft, a 100-person Indian thermoforming machinery company, built Ira to preserve knowledge that had previously lived in the heads of three... - [Should AI Engineers Still Read Code in 2026? The Z/L Continuum](https://aietalks.com/talks/should-ai-engineers-still-read-code-in-2026-the-z-l-continuum): Alex Volkov, ThursdAI. Alex Volkov frames the argument over AI-generated code through two opposing conference messages. Ryan Lopopolo says code is free and engineers should move their attention... - [Understanding is the new bottleneck](https://aietalks.com/talks/understanding-is-the-new-bottleneck): Geoffrey Litt, Notion. Geoffrey Litt argues that human understanding remains necessary as coding agents produce larger and more complex changes. Correctness checking is becoming easier to automate, but... - [The Golden Age of AI Engineering](https://aietalks.com/talks/the-golden-age-of-ai-engineering): Alexander Embiricos, Romain Huet & Peter Steinberger, OpenAI. Alexander Embiricos and Romain Huet describe how quickly AI engineering has changed, from models that generated code without testing it to agents that can pursue... - [Building an ACP-Compatible Agent Live](https://aietalks.com/talks/building-an-acp-compatible-agent-live): Bennet Fenner, Zed. Bennet Fenner builds a small TypeScript coding agent and connects it to Zed through the Agent Client Protocol. He first explains ACP as a JSON-RPC... - [Everything We Knew About Software Has Changed](https://aietalks.com/talks/everything-we-knew-about-software-has-changed): Theo Browne, @t3dotgg. Theo Browne argues that software engineers are still building around assumptions from an earlier era. He describes Sonnet 3.5 as the reliable tool-calling era, Opus... - [I Run a Fleet of AI Agents Across Three Machines. Here's What Broke.](https://aietalks.com/talks/i-run-a-fleet-of-ai-agents-across-three-machines-heres-what-broke): Kyle Jaejun Lee, KRAFTON. Kyle Jaejun Lee describes running AI coding agents across a MacBook and two always-on Linux machines. His first problem was human attention: six flat agent... - [Teaching Coding Agents to Do Spreadsheets](https://aietalks.com/talks/teaching-coding-agents-to-do-spreadsheets): Nuno Campos, Witan Labs. Nuno Campos describes four months of work teaching coding agents to handle spreadsheets. The team moved from about 50% to 92% accuracy on an internal... - [Think You Can Build a Game with AI? Think Again!](https://aietalks.com/talks/think-you-can-build-a-game-with-ai-think-again): Danielle An & David Hoe, Meta. Danielle An and David Hoe describe how AI is changing game creation at Meta. Basic games such as platformers and Tetris can now be generated... - [Your agent is blindfolded](https://aietalks.com/talks/your-agent-is-blindfolded): Johan Lajili, Poolside AI. Johan Lajili argues that the difference between successful and disappointing agent use is often the feedback loop, rather than whether the project is greenfield or... - [Your coding agent doesn't always follow your rules](https://aietalks.com/talks/your-coding-agent-doesnt-always-follow-your-rules): Talha Sheikh, Checkout.com. Talha Sheikh describes the gap between an agent completing a coding task and completing it according to the developer's requirements. His first response was Vector... - [Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data](https://aietalks.com/talks/your-llm-deception-monitor-is-broken-the-fix-is-in-the-training-data): Sachin Kumar, LexisNexis. Sachin Kumar argues that behavioral tests cannot reliably detect sleeper-agent backdoors because the model behaves normally until it sees a specific trigger. He also examines... - [500 People Vibe-Coded for 30 Days. I Was One of Them.](https://aietalks.com/talks/500-people-vibe-coded-for-30-days-i-was-one-of-them): Sanja Grbic, Automattic. Sanja Grbic describes Automattic's Radical Speed Month, a 30-day experiment in which about 500 employees started around 794 projects after pausing roadmap work. She built... - [Beyond the Harness: A Journey Towards Adaptive Engineering](https://aietalks.com/talks/beyond-the-harness-a-journey-towards-adaptive-engineering): Rajiv Chandegra, Annicha Labs. Rajiv Chandegra argues that current AI engineering depends on fixed harnesses. These harnesses define roles, tools, prompts, sequencing, memory, and handoffs before a run starts.... - [Build AI Systems for Discernment, Not Approval](https://aietalks.com/talks/build-ai-systems-for-discernment-not-approval): Angel Ortmann Lee, Duolingo. Angel Ortmann Lee argues that adding a human to an AI workflow does not guarantee careful judgment. People often accept AI output with little scrutiny,... - [GTM Is You](https://aietalks.com/talks/gtm-is-you): Victoria Melnikova, Evil Martians. Victoria Melnikova argues that distribution has become the main constraint for developer tool companies because building software is easier and AI has filled feeds and... - [How We Taught Agents to Use Good Retrieval](https://aietalks.com/talks/how-we-taught-agents-to-use-good-retrieval): Hanna Lichtenberg, Mixedbread AI. Hanna Lichtenberg argues that knowledge agents are often limited by retrieval rather than reasoning. On BrowseComp Plus and Office QA Pro, models perform far below... - [Respect The Process](https://aietalks.com/talks/respect-the-process): Andrew Dumit, Watershed Technology Inc.. Andrew Dumit describes an agent that edits large supply-chain graphs used to build product carbon footprints. Early versions relied on specialized function calls, but they... - [SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale](https://aietalks.com/talks/swe-marathon-evaluating-coding-agents-at-billion-token-scale): Rishi Desai, Abundant AI. Rishi Desai presents SWE-Marathon, a benchmark for autonomous software work that runs far beyond individual bug fixes. Its 20 tasks cover library rewrites, full-stack product... - [The Pipeline Is Dead](https://aietalks.com/talks/the-pipeline-is-dead): Iris ten Teije, Sky Valley Ambient Computing. Iris ten Teije argues that software distribution still assumes one frozen version for everyone because changing software used to require expensive, central work. That constraint... - [Field Guide to Fable](https://aietalks.com/talks/field-guide-to-fable): Thariq Shihipar, Anthropic. Thariq Shihipar presents Fable as a model whose capabilities appear in uneven spikes. A chat model may know every Pokemon name yet fail to find... - [Continual Learning for AI Agents: From Failures to Durable Improvements](https://aietalks.com/talks/continual-learning-for-ai-agents-from-failures-to-durable-improvements): Soheil Feizi, RELAI. Soheil Feizi presents continual learning as a way for AI agents to improve from real use without forgetting prior capabilities. He separates the agent into... - [MCP Apps: Primitives, Discovery, and the Future of Software](https://aietalks.com/talks/mcp-apps-primitives-discovery-and-the-future-of-software): Pietro Zullo, Manufact, Inc. Pietro Zullo explains MCP Apps as an interaction layer for MCP servers. A server can return a UI resource declared with a `ui://` resource, which... - [The Missing Layer After Launch](https://aietalks.com/talks/the-missing-layer-after-launch): Raphael Kalandadze, Wandero AI. Raphael Kalandadze argues that launching an agent product is only the beginning. Agent failures often appear inside successful-looking conversations: a wrong price, a dropped constraint,... - [Your AI Product Will Fail Unless You Can Explain It](https://aietalks.com/talks/your-ai-product-will-fail-unless-you-can-explain-it): Veronica Hylak, Hey AI. Veronica Hylak presents a three-part method for explaining complex AI products. First, start with the customer's wound: the frustrating moment in their work, such as... - [The Prompt Is Still a Punch Card](https://aietalks.com/talks/the-prompt-is-still-a-punch-card): Ted Johnson, JoinIn AI. Ted Johnson argues that AI has improved what people can express to computers without changing the basic interaction protocol. A prompt is still a packaged... - [Software Factories & Keynotes](https://aietalks.com/talks/software-factories-keynotes): Swix, AI Engineer & Pablo Castro, Microsoft & Alexander Emiricos, Roman Huitt & Peter Steinberger, OpenAI & Tashan, Z.ai & Thomas Wolf, Hugging Face & Olive, MiniMax & Randall Daggs, Snyk & Theresa, Factory & Kushan, independent & Simon Eskildsen, Turbopuffer & Gergely Orosz, The Pragmatic Engineer & Zach Lloyd, Warp & Gabe, OpenGov & Solomon Hykes, Dagger & Kyle, HumanLayer & Dominic Tornow, Resonate & Eric Meyer, Leantime's Labs & Lee Robinson, Cursor & Sarah, Notion & Vibhor Gupta, BAML & Jack, HumanLayer. This day of keynote programming examines what it takes to build software factories around coding and general-purpose agents. Speakers describe systems that collect work, plan... - [Building Great Agent Skills: The Missing Manual](https://aietalks.com/talks/building-great-agent-skills-the-missing-manual): Matt Pocock, AI Hero. Matt Pocock presents a checklist for escaping what he calls "skill hell", the problem of having many available agent skills without knowing which ones work... - [Deterministic Infra for Non-Deterministic AI Agents](https://aietalks.com/talks/deterministic-infra-for-non-deterministic-ai-agents): Nishant Gupta, Meta Superintelligence Labs. Nishant Gupta argues that autonomous agents should be treated as distributed systems. Traditional cloud infrastructure assumes short-lived requests, known execution paths, bounded failures, and mostly... - [Frontier Results, On Device](https://aietalks.com/talks/frontier-results-on-device): RL Nabors, Arize. RL Nabors argues that many calls sent to frontier models could run on smaller language models or task-specific models. Local inference can keep data on... - [The Agentic AI Engineer](https://aietalks.com/talks/the-agentic-ai-engineer): Benedikt Sanftl & Burak, Mutagent. Benedikt Sanftl and Burak describe an Agentic AI engineer as a multi-agent system that builds and improves AI agents through two connected loops. The offline... - [The Future Is Domain-Specific Agents](https://aietalks.com/talks/the-future-is-domain-specific-agents): Justin Schroeder, StandardAgents. Justin Schroeder argues that useful AI systems should be built from many focused agents rather than one general agent loaded with tools, skills, and context.... - [The Prompt is the Platform](https://aietalks.com/talks/the-prompt-is-the-platform): Dominik Tornow, Resonate HQ, Inc. Dominik Tornow argues that coding agents will reduce the value of general-purpose software platforms by generating bespoke implementations for the infrastructure already in place. Resonate's... - [You Can't Prompt the Room: The Last Skill AI Won't Replace](https://aietalks.com/talks/you-cant-prompt-the-room-the-last-skill-ai-wont-replace): Balázs Horváth, VisualLabs. Balázs Horváth argues that AI has made implementation cheap while making discovery and decision-making more valuable. VisualLabs generated 21 agent ideas in an internal hackathon,... - [Your Agent Failed in Prod. Good Luck Reproducing It.](https://aietalks.com/talks/your-agent-failed-in-prod-good-luck-reproducing-it): Tisha Chawla & Susheem Koul, Microsoft. Tisha Chawla and Susheem Koul argue that production agents need replayability rather than bitwise deterministic model output. Setting temperature to zero does not remove variation... - [Agents Building Agents](https://aietalks.com/talks/agents-building-agents): Alfonso Graziano, Nearform. Alfonso Graziano describes building production AI agents with a second AI system that writes and changes their code. The coding agent improves a product agent... - [AI-Driven Multi-Document Correlation for Financial Compliance](https://aietalks.com/talks/ai-driven-multi-document-correlation-for-financial-compliance): Varsha Shah, Independent Researcher. Varsha Shah presents a framework for finding compliance risks that individual document checks miss. It connects employees, vendors, accounts, transactions, and regulatory records in a... - [AI System Design: From Idea to Production](https://aietalks.com/talks/ai-system-design-from-idea-to-production): Apoorva Joshi, MongoDB. Apoorva Joshi presents a repeatable framework for designing AI systems from an initial business problem through production. She applies it to a health insurance claims... - [Browser Agents Don't Need Better Models. They Need Better Eyes.](https://aietalks.com/talks/browser-agents-dont-need-better-models-they-need-better-eyes): Kushan Raj, ARK. Kushan Raj argues that browser-agent progress has focused too much on upgrading models. In his examples, agents using screenshots and ordinary browser interaction spend a... - [Building an Autonomous Engineering Org](https://aietalks.com/talks/building-an-autonomous-engineering-org): Angie Jones, Agentic AI Foundation. Angie Jones describes Block's attempt to move a 3,500-person engineering organization from occasional AI assistance to agent-led software delivery. Early adoption was high, but delivery... - [Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry](https://aietalks.com/talks/bypassing-the-multimodal-tax-hybrid-rag-sql-rrf-ui-telemetry): Abed Matini, Ogilvy. Abed Matini presents a local-first FAQ chatbot for an employee handbook. The application converts PDFs, Word files, PowerPoint files, and images into Markdown before chunking... - [HTML is All You Need (for Agents to Make Graphics)](https://aietalks.com/talks/html-is-all-you-need-for-agents-to-make-graphics): Amol Kapoor, Nori. Amol Kapoor argues that coding agents are poor at graphics mainly because people give them tools designed for human hands and eyes. Canvas tools such... - [OpenClaw in Your Hand: Building a Physical AI Terminal](https://aietalks.com/talks/openclaw-in-your-hand-building-a-physical-ai-terminal): Lech Kalinowski, Callstack. Lech Kalinowski built Vault as a physical remote control for an OpenClaw instance running on an NVIDIA DGX Spark. The handheld uses an ESP32 dual-core... - [Research to Reality: Bringing Frontier ML Research to Production](https://aietalks.com/talks/research-to-reality-bringing-frontier-ml-research-to-production): Vaidas Razgaitis, Higharc. Vaidas Razgaitis describes the handoff between ML researchers who build novel prototypes and software engineers who turn those prototypes into production features. He treats the... - [Structuring the Unstructured](https://aietalks.com/talks/structuring-the-unstructured): Cedric Clyburn, Red Hat. Cedric Clyburn argues that document processing determines whether an AI application can use enterprise data accurately. PDFs often contain tables, images, diagrams, headers, and scanned... - [The 100-Tool Agent Is a Trap](https://aietalks.com/talks/the-100-tool-agent-is-a-trap): Sohail Shaikh & Ankush Rastogi, Prosodica. Sohail Shaikh and Ankush Rastogi describe why the common fat-agent design degrades as an agent gains tools. Every request carries every function name, description, and... - [User Signal Dies at the Retrieval Boundary](https://aietalks.com/talks/user-signal-dies-at-the-retrieval-boundary): Sonam Pankaj, StarlightSearch Inc. Sonam Pankaj argues that production agents need to learn from task outcomes at runtime. Current systems collect tool calls, model completions, exceptions, and pass or... - [Using Spec-Driven Development for Production Workflows](https://aietalks.com/talks/using-spec-driven-development-for-production-workflows): Erik Hanchett, Amazon Web Services. Erik Hanchett presents spec-driven development as a way to control coding assistants during larger software changes. The workflow starts with written requirements, moves to a... - [Voice In, Visuals Out: The Agony and the Ecstasy](https://aietalks.com/talks/voice-in-visuals-out-the-agony-and-the-ecstasy): Allen Pike, Forestwalk Labs. Allen Pike argues that voice is a strong input method because people speak faster than they type and communicate meaning through tone, while visuals are... - [We Cut 94% of AI Coding Tokens With a Local Code Index](https://aietalks.com/talks/we-cut-94-of-ai-coding-tokens-with-a-local-code-index): Rajkumar Sakthivel, Tesco. Rajkumar Sakthivel describes how his team responded to a sudden increase in its AI coding bill. Their investigation found that coding tools were sending about... - [When All Context Matters: Extended Cache Augmented Generation](https://aietalks.com/talks/when-all-context-matters-extended-cache-augmented-generation): Luis Romero-Sevilla, Orbis Operations. Luis Romero-Sevilla describes a document collection where every document contributes to answering a question, while the documents are replaced frequently. Simple RAG can refresh its... - [Your Agent Is Wasting Tokens and You Don't Know It](https://aietalks.com/talks/your-agent-is-wasting-tokens-and-you-dont-know-it): Erik Hanchett, Amazon Web Services. Erik Hanchett describes five ways to reduce token costs in production agents without changing prompts or switching away from the models already in use. First,... - [A Genius With Amnesia](https://aietalks.com/talks/a-genius-with-amnesia): Victor Savkin, Nx. Victor Savkin compares current coding agents with a brilliant engineer who can see only a tiny part of the codebase and forgets every conversation. In... - [Agents in Production: How OpenGov Built and Scaled OG Assist](https://aietalks.com/talks/agents-in-production-how-opengov-built-and-scaled-og-assist): Gabe De Mesa, OpenGov. Gabe De Mesa explains how OpenGov built OG Assist, an agent available across its ERP products for budgeting, procurement, asset management, permitting, and other government... - [Stop Writing Tone Instructions. Layer Them.](https://aietalks.com/talks/stop-writing-tone-instructions-layer-them): Isadora Martin-Dye, Isadora & Co | The Bloom House AI. Isadora Martin-Dye argues that brand voice is an architecture rather than a tone instruction added to one system prompt. Her four layers are immutable identity... - [Turn 10,994 Notes Into Memory](https://aietalks.com/talks/turn-10-994-notes-into-memory): Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI. Paul Iusztin and Louis-François Bouchard show how they turned years of notes, videos, documents, and code repositories into a personal AI research OS. Their system... - [Build Systems, Not Code](https://aietalks.com/talks/build-systems-not-code): Angie Jones, Agentic AI Foundation. Angie Jones argues that engineers do not lose the satisfaction of building when coding agents write more of the implementation. The work moves to designing... - [Production Evals For Agentic AI Systems](https://aietalks.com/talks/production-evals-for-agentic-ai-systems): Nishant Gupta, Meta Superintelligence Labs. Nishant Gupta argues that benchmark scores do not tell teams whether an autonomous system will behave reliably in production. Agents plan, call tools, retrieve information,... - [Recursive Coding Agents](https://aietalks.com/talks/recursive-coding-agents): Raymond Weitekamp, OpenProse. Raymond Weitekamp argues that coding agents are currently "mismanaged geniuses." Their raw intelligence is often sufficient, but they do not reliably deliver outcomes. He presents... - [The Log Is The Agent](https://aietalks.com/talks/the-log-is-the-agent): Ishaan Sehgal, Omnara. Ishaan Sehgal argues that an agent should be understood as its durable session log rather than its model, runtime, tools, or worker process. The log... - [The Miranda Hypothesis: How Hamilton Poisoned Persona Evals](https://aietalks.com/talks/the-miranda-hypothesis-how-hamilton-poisoned-persona-evals): Jacob E. Thomas, Results Gen. Jacob E. Thomas argues that persona systems often produce culturally familiar composites instead of historically grounded people. A model can sound like Alexander Hamilton while... - [6 Things to Know about AIE World's Fair 2026](https://aietalks.com/talks/6-things-to-know-about-aie-worlds-fair-2026): Swix, AI Engineer. Swix gives a practical tour of what is new at AIE World's Fair 2026. The event is larger than previous AIE gatherings, has an extra... - [The Production AI Playbook: Deploying Agents at Enterprise Scale](https://aietalks.com/talks/the-production-ai-playbook-deploying-agents-at-enterprise-scale): Sandipan Bhaumik, Databricks. Sandipan Bhaumik presents a five-pillar framework for moving AI agents from demos into production: evaluation, observability, data foundation, orchestration, and governance. His central advice is... - [Your Agent's Biggest Lie: "I Searched the Web"](https://aietalks.com/talks/your-agents-biggest-lie-i-searched-the-web): Rafael Levi, Bright Data. Rafael Levi describes a failure mode where an agent cannot access a website but hides that failure behind a confident answer. Anti-bot systems can return... - [You Might Not Need 50 Diffusion Steps](https://aietalks.com/talks/you-might-not-need-50-diffusion-steps): Ziv Ilan, NVIDIA. Ziv Ilan explains how NVIDIA is making image and video diffusion practical for developers and enterprises. Diffusion models often need 20 to 50 denoising steps,... - [Why MCP and ChatGPT Apps Use Double Iframes](https://aietalks.com/talks/why-mcp-and-chatgpt-apps-use-double-iframes): Frédéric Barthelet, Alpic. Frédéric Barthelet explains why ChatGPT renders MCP app views through two nested iframes. A direct `srcdoc` iframe shares ChatGPT's origin, so ChatGPT's content security policy... - [The agent-ready web: Simplify user actions with WebMCP](https://aietalks.com/talks/the-agent-ready-web-simplify-user-actions-with-webmcp): Tara Agyemang, Google Chrome. Tara Agyemang introduces WebMCP, a proposed web standard for exposing website capabilities as structured tools to in-browser AI agents. She contrasts direct tool calls with... - [Why Can't Anyone Answer Questions About the Business?](https://aietalks.com/talks/why-cant-anyone-answer-questions-about-the-business): Garrett Galow, WorkOS. Garrett Galow describes Studio, an internal WorkOS workspace for answering business questions and building reusable tools. The problem is familiar: someone asks for data, explains... - [Your Attention Is the Bottleneck, Not Your Agents](https://aietalks.com/talks/your-attention-is-the-bottleneck-not-your-agents): Zack Proser, WorkOS. Zack Proser argues that AI coding tools have shifted the limiting factor from agent capacity to human attention. He describes a Slack bug fix where... - [Self Driving Products: Product Signals to Pull Requests](https://aietalks.com/talks/self-driving-products-product-signals-to-pull-requests): Joshua Snyder, PostHog. Joshua Snyder describes PostHog's attempt to turn observability data into pull requests. The system ingests signals from product analytics, error tracking, logs, Slack, and session... - [Sovereign Escape Velocity: Ownership with Open Models](https://aietalks.com/talks/sovereign-escape-velocity-ownership-with-open-models): Gus Martins & Ian Ballantyne, Google DeepMind. Gus Martins and Ian Ballantyne explain why an organization might choose an open model even when a stronger hosted model is available. Gemma 4 is... - [Stop Making Models Bigger, Make Them Behave](https://aietalks.com/talks/stop-making-models-bigger-make-them-behave): Kobie Crawford, Snorkel. Kobie Crawford presents a study from Snorkel and UC Berkeley's RLLM team on improving financial analysis with reinforcement learning. A 235 billion parameter Qwen 3... - [2026 AI Engineer Vibe Reel](https://aietalks.com/talks/2026-ai-engineer-vibe-reel): . This is a very short promotional reel built around repetition rather than explanation. The spoken audio consists of the word "Heat" repeated several times, with... - [GPU Cloud Deployment Without Leaving Your IDE](https://aietalks.com/talks/gpu-cloud-deployment-without-leaving-your-ide): Audrey Hsu, RunPod. Audrey Hsu explains how RunPod's Flash Python SDK shortens the development loop for GPU inference. The usual process requires committing code, building a Docker image,... - [RAG is dead, right??](https://aietalks.com/talks/rag-is-dead-right): Kuba Rogut, Turbopuffer. Kuba Rogut argues that the claim that RAG is dead comes from defining RAG too narrowly. A retrieval system can combine semantic search with full-text... - [Road to 5 Million Tokens: Breaking Barriers in Long Context Training](https://aietalks.com/talks/road-to-5-million-tokens-breaking-barriers-in-long-context-training): Max Ryabinin, Together AI. Max Ryabinin explains how Together AI fit extremely long-context training onto a single 8xH100 node. A standard Llama 3B model cannot fit even its parameters... - [Why Eval++ Is the Next Great Compute Primitive](https://aietalks.com/talks/why-eval-is-the-next-great-compute-primitive): Sunil Pai & Matt Carey, Cloudflare. Sunil Pai and Matt Carey explain why Cloudflare uses Durable Objects as the foundation for AI agents. A Durable Object gives each identifier a persistent... - [Why More Context Makes Your Agent Dumber and What to Do About It](https://aietalks.com/talks/why-more-context-makes-your-agent-dumber-and-what-to-do-about-it): Nupur Sharma, Qodo. Nupur Sharma describes how agent systems fail as their context and orchestration become more complicated. Models may accept large inputs while still losing information from... - [From MCP to Scale: Pipelines That Build Themselves](https://aietalks.com/talks/from-mcp-to-scale-pipelines-that-build-themselves): Rafael Levi, Bright Data. Rafael Levi argues that the expensive part of web data work is maintaining scrapers, especially on sites with changing selectors, React interfaces, or strong anti-bot... - [LLM Observability, Evaluation, Experimentation Platform](https://aietalks.com/talks/llm-observability-evaluation-experimentation-platform): Dat Ngo, Arize AI. Dat Ngo describes AI application development as software engineering with nondeterministic execution paths. Observability begins with OpenTelemetry traces and spans, which record an agent's actions,... - [Under 5 minutes to a deployed LLM endpoint](https://aietalks.com/talks/under-5-minutes-to-a-deployed-llm-endpoint): Audry Hsu, RunPod. Audry Hsu introduces RunPod as a cloud AI infrastructure company that provides GPUs and deployment tools for private models and open-source models from Hugging Face.... - [Building Interactive UIs in VS Code with MCP Apps](https://aietalks.com/talks/building-interactive-uis-in-vs-code-with-mcp-apps): Marlene Mhangami & Liam Hampton, Microsoft and GitHub. Marlene Mhangami and Liam Hampton explain how MCP Apps add interactive interfaces to MCP tool results in VS Code. MCP standardizes how applications provide context... - [Building Safe Payment Infrastructure for the Autonomous Economy](https://aietalks.com/talks/building-safe-payment-infrastructure-for-the-autonomous-economy): Steve Kaliski, Stripe. Steve Kaliski argues that agents already spend money through model subscriptions and cloud tools, even when they do not directly pay ordinary businesses. The risk... - [Evals Are Broken, Use Them Anyway](https://aietalks.com/talks/evals-are-broken-use-them-anyway): Ara Khan, Cline. Ara Khan argues that evals fail when engineers treat benchmark scores as objective truth or replace measurement with personal preference. Scores can still guide engineering... - [Beyond Transcription: Building Voice AI That Understands Conversations](https://aietalks.com/talks/beyond-transcription-building-voice-ai-that-understands-conversations): Hervé Bredin, pyannoteAI. Hervé Bredin argues that speech-to-text alone leaves out information needed to understand conversations. A useful system needs to identify who speaks when, attach speakers to... - [Building Agent Interfaces: Lessons from Chrome DevTools (MCP) for Agents](https://aietalks.com/talks/building-agent-interfaces-lessons-from-chrome-devtools-mcp-for-agents): Michael Hablich, Google. Michael Hablich explains how the Chrome DevTools team built an interface for coding agents and had to redesign it several times. Their first approach exposed... - [The Art & Science of Benchmarking Agents](https://aietalks.com/talks/the-art-science-of-benchmarking-agents): Vincent Chen, Snorkel AI. Vincent Chen argues that benchmarks have fallen behind agent capabilities. Enterprises may be impressed by coding agents, yet still hesitate to deploy them in high-stakes... - [Software Fundamentals Matter More Than Ever](https://aietalks.com/talks/software-fundamentals-matter-more-than-ever): Matt Pocock. Matt Pocock argues that AI has made software fundamentals more important because bad code can now accumulate faster. He tried a specs-to-code workflow that repeatedly... - [Don't Build Agents, Build Skills Instead](https://aietalks.com/talks/dont-build-agents-build-skills-instead): Barry Zhang & Mahesh Murag, Anthropic. Barry Zhang and Mahesh Murag argue that many teams should stop creating separate agents for every domain. A general agent with a runtime, file system,... - [No Vibes Allowed: Solving Hard Problems in Complex Codebases](https://aietalks.com/talks/no-vibes-allowed-solving-hard-problems-in-complex-codebases): Dex Horthy, HumanLayer. Dex Horthy argues that coding agents struggle with established codebases because their context becomes noisy, incomplete, or pointed in the wrong direction. His answer is... ## Index pages - [Talks from 2026](https://aietalks.com/talks/2026): 348 talks - [Talks from 2025](https://aietalks.com/talks/2025): 2 talks - [All packs](https://aietalks.com/packs) - [By conference](https://aietalks.com/conferences) - [By speaker](https://aietalks.com/speakers) - [By company](https://aietalks.com/companies) - [By topic](https://aietalks.com/tags) - [RSS feed](https://aietalks.com/rss.xml)