Talks
348 talks in 2026
Agent Spending Without Controls
Beyond the Lethal Trifecta: Agentic Commerce on the Open Internet
Multimodal Collaborative Agents for Next-Gen Commerce
Teaching agents to pay
The End of the Static Screen: Architecting Intent-Driven UX
When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS
Why Your AI Agent Needs a Wallet: USDC and Nanopayments
x402 Isn't Good (Yet)
Your Agent Just Authorized What?!
SOTA Generative Media Panel
Agentic Sites: Building Hyper Personalized Websites
Agents Are Where Microservices Were in 2015
AI Agents Are Just Distributed Systems Now
From Tokenmaxxing to Trusted Throughput
Tell the Robot What You Want
The Half Life of Agent Infrastructure
The Signal Layer: What to Build When Anything Can Be Built
Tribal Dungeons of Global Shipping: AI Agents at Global Scale
Which AI startups actually land enterprise contracts?
AI Evals for Cross-Functional Teams
AI-Native Organisations Run on Skills: How to Structure and Scale Them
Building the Engine While Flying the Plane: Launching the Figma MCP Server
Building uReview, Uber's Multi-Agent Code Review Engine
From AI-Assisted to AI-Native: Building a Frontier Development Team
How do you diffuse AI into the real world?
How to Avoid Disaster When Vibe-Coding a Billing Engine
How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage)
Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons
Your Code Has Bugs. Lean4 Has Proofs: Formal Verification for Engineers
Can LLMs Write Fast Multi-GPU Kernels?
How Anthropic Builds: Lessons from Labs
How to Generate Mergeable Code with a Context Engine
KV Cache-Aware Routing and P/D Disaggregation on Kubernetes
The Agentic Commerce Stack
AI in GTM at Notion
Building GTM AI Agents: Lessons from Deploying to 6,000 Users
GTM Engineering: The Technical Bits
How AI Agents Let GTM Teams Scale
How We Got LLMs to Recommend Our Open Source Library
Knowledge Systems: The New GTM Stack
Reverse-Engineering the AI Buyer
The Building Blocks of GTM Orchestration
The Death of Developer Advocates
The Missing Layer in Agentic AI
Einstein Arena: Harnessing Collective Agent Intelligence for Open Science
Agent Frameworks Considered Harmful
FinOps for AI Agents: Who Spent All the Tokens?
Give the Agent a Budget, Not a Token
Preferences Over Benchmarks: Model Routing
The Agent Behind the Curtain: Building the Oz Cloud Agent Platform
What If Your Chip Design Team Moved Like a Single Body?
Agentic SDLC at Uber
Building Agents Is Trivial Now, Context Is the Next Frontier
The Missing Layer: Design Taste in AI Agents
How I automate my own job at Hugging Face using agents
IT Admin for the AI Workforce
Prototyping as Leadership: How a CTO Ships with AI Agents
The Era of Compound Engineering
The Last Human Code Review: Building Trust in AI-Generated Code
Unlock Agent Autonomy: The Runtime for AI-Native Systems
Your Agent Evolved. Your Evals Didn't.
Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards
200 Million Patient Interactions Later
AI is the World's Largest Relationship Therapist
Don't Be Data Poor
From Ambient Documentation to Clinical Intelligence
Guardrails First: Engineering Member-Facing Health AI
Healthcare's Agent Bytecode: X12 as the Harness for AI Agents
How to build an AI-Native Health Company
Shipping AI to a Million Patients Without an A/B Test
Trading Desks to Clinical Trials: Parallels in Applied Vertical AI
Why Your Enterprise Tech Stack Isn't Ready for AI Agents
Building an Agentic Video Editor for Mass Consumer
Generative Video at the Speed of Light
Infra behind Krea 2: How to train and serve at scale
The Next Game Engine Won't Have a Manual
The Next Medium: Why Real-Time Interactive Video Changes Everything
Training Krea 2: What Matters in Generative Model Training
Voice agents with Realtime Video
While My Guitar Gently Speaks
Context Engineering in 2026
How to Kill the Code Review
Security Firewall for Agents
Bringing agents onto the World Wide Web
Computer Use at the Edge of the Statistical Precipice
Computer-use models will agentify the web, not APIs
From RL to IRL
How Web Data Infrastructure Powers the Next Generation of AI
The Dark Arts of Web Automation: Teaching Agents to Use Websites Like Humans
The Rise of CaaS: Context-as-a-Service for Agentic AI
Beyond Static Intelligence: Evaluating Continual Learning
Bringing Continual Learning into Enterprises
Designing Agents (The Floor Is the Frontier)
Gradient-Free Continual Learning
Improving Agents is a Data Mining Problem
Intelligence + Continual Learning = Expertise
Lessons from Studying Every Memory System
LLM Knowledge Bases: A Practical Guide
Memory Harnesses for Long-Running Research Agents
Scaling Compute on Context
Scaling up Continual Learning
Agents, codebases, and teams
The Evolution of Agentic Surfaces
Codex, Behind the Harness
Taking Reinforcement Learning Cross Datacenter
Always-on agents run production without the on-call tax
Guide, Verify, Solve
Multiplayer agentic engineering
Velocity Sickness: What Happens When Your Whole Team Gets 10x Faster
Anthropic's CCA Exam as a Field-Guide for Agentic Engineering
Benchmarking Coding Agents on New vs Legacy Codebases
Realtime multiplayer, automation, and you!
Compression at the Edge
Local Models: Trust, Control, Optimization
Open Source Is Dead. Long Live Open Source.
The New Primitives: Building AI Native Software
The State of Model Routing
Gadgets: Personal app vibe coding that is actually safe
Building Turbopuffer
MCP Apps: Extending the Frontier
MCP Tasks (async): Why Aren't Any Agents Supporting Them?
When Will the Benchmaxxing Plague End?
Rethinking Environments for Long-Horizon Work
Teaching AI to Find Real Vulnerabilities
Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It
Benchmarks: The Good, the Bad, and the Ugly
Data and Environment Curation for Post-Training LLMs
Data Quality Is the Compute Multiplier
Ending AI Slop
Fighting Slop with Slop
Learning on the Job: The Future of Post-Training
Reinforcement Learning without Verifiable Rewards
Scaling to Long Horizons
The Base Model Is Dead
The Data for Fully Autonomous Software Engineers and Companies
Verifiable Environments for AI in Biology
What's Next After RLHF?
Build for the Memo, Not the Demo
First Steps Toward Automated AI Research
Let's integrate AI Agents in Event-Sourced Systems
Your Finance Agent's Bottleneck Is You
How Kepler Built Verifiable AI for Financial Services
Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains
Persona Engineering: A Field Guide to AI Synthetic Personas
SimulationMaxxing: How We Ship Agents 20× Faster
Skills are New Features: Building a Skill-Centric Harness
We Vetted 2,000 AI Skills Before They Reached Developers
Wearing the Agent: From Group Chats to Glasses
Why Off-the-Shelf AI Doesn't Understand Money
Your Agent Didn't Fail. Your Harness Did.
AI Agents for Performance: Ship Faster, Pay Less
AI tools for Forward Deployed Engineering
Forward Deployed Engineering 101
How Forward Deployed Engineering is Done at Cognition
How Forward Deployed Engineering Is Done at Decagon
How Forward Deployed Engineering is done at Factory
How Forward Deployed Engineering is done at Kepler
How Forward Deployed Engineering is done at Ramp
Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub
The Dirty Secret of Forward Deployed Engineering
DeepSWE: A Contamination-Resistant Coding Benchmark
State of Data
The Messy Reality of Scale: Synthetic Data and Pre-Training
Evaling Video Slop
Evals-Driven Development for a Mental Health AI Coach
From Agent Traces to Agent Simulations
Loop Engineering from First Principles
Why Large? Tiny LMs and Agents on Edge and Robotics
Building Closed-Loop Evals for a Multimodal Agent at Scale
Everything Is a Rollout
From Signal to PR: Anatomy of a Self-Improving Agent
How Evals and Prompts Shape Agent Behavior
Setting Yourself Up for Success
The Future of Evals: From LLM as a Judge to Agent as a Judge
Training Frontier Models to Out-Think Hackers
Vending-Bench: Long-Horizon Agent Evals
AI on Your Lakehouse: Context Comes in Shapes, Not Queries
Citation Needed: Provenance for LLM-Built Knowledge Graphs
Harness Engineering is Not Enough: Why Software Factories Fail
Learned Execution Graphs for Anomaly Detection & Drift in APIs
Local Agentic Theory For Mobile Games
Notion's Token Town
Perception Agents
The Unreasonable Effectiveness of Separating the Task from the Model
Video Has No Memory. Here's How We Built One.
Why Agentic Systems Need Ontologies
Why We Killed Our Multi-Agent Pipeline
Active Graph Agent Runtime (BabyAGI 4)
Claude for Long-Horizon Tasks
CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens
From Systems of Record to Systems of Context
Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer
Your Moat Is Your Data Model
Better Auth
Every Harness Will Become A Claw
HTML Is All Agents Need
The Biggest Challenge in Your Stack? Evals, Evals, Evals
The Desktop Frontier
Your agent architecture has a half-life of 6 months
Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing
Build the AI GTM Agent That Knows the Buyer
Can Oncology Workflows Run Without Human Touch?
Designing Voice Agents for Real Conversations
Don't Let the LLM Drive
Enterprise Agents Have a Structure Problem
In the Land of AI Agents, the Verifiers Are King
It's 10pm. Do You Know Where Your Agents Are?
Medic for Apache Spark: First Aid for Failing Jobs
Security Track Intro
Skills are the New SDKs
We Gave an Agent Production Code Access and Then Tried to Sleep at Night
When Agents Meet Physical Data: The Other Physics of Agent Harnesses
Why Your Agent Disagrees With Itself (And What To Do About It)
Your LLM Stack Is a 2008 Database With Better Marketing
Your Voice Agent Doesn't Need a Frontier Model
Build Evals That Actually Matter
From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization
From Tokens to Cells: Foundation Models for Single-Cell Biology
You Didn't Ship a Bug. You Just Wrote It for a Human.
A Practitioner's Guide to Graphs
Agents Need Feature Flags
Agents Need Receipts, Not More Tool Calls
Autonomous Agents for Scientific Tasks
Content Is Code
Stop Burning Tokens: Why Self-Improvement Needs Domain Expertise First
Stop Renting Your Cognitive Infrastructure
The UX of AI: Making AI-Powered Apps Your Users Don't Hate
Your Agents Need a Save Button
Every company should have a Brain
On AI and Knowledge
Software engineering is not about writing code
Special Topics in Kernels, RL, Reward Hacking in Agents
The Great Loops Debate
Using LLMs to Secure Source Code
Imagination Engineering: "Live in the future and then build what's missing."
Computer-Use 2.0: Agents Just Got Multi-Cursor
Recursive Model Improvement
Don't Ship Skills Without Evals
Forward Deployed Engineering at Cursor
The engineer of the future is the person who is able to choose what is worth doing.
WTF Is the Context Layer? The Missing Infrastructure for Production Agents
From fork() to Fleet: Designing an Agent Sandbox Cloud
I've Never Seen Anything Scarier Than an LLM with Tool Calls
Modern Post-Training: A Deep Dive
Stop Evaluating Models Like It's the 50s
A Song of Types and Agents
remobi.app: Don't Change Your Terminal Workflow for Mobile
ReviewDebt: A Practical Framework for Scoring Every Pull Request
Semantic Blindness: 500,000 Sensors Confused an LLM
The Agentic Web and the Bazaar Era of AI
The AI bugpocalypse is here. Now what?
What Does Done Even Mean? Agents and Paperclip's Liveness Model
Chat and citations won't save your vertical AI
Claws Out: Securing and Building with OpenClaw
Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers
Develop at Idea Velocity
Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD
From Writing Code to Designing Systems: How the Developer Role is Changing
State of the Union: Why Local, Why Now
Stop AI Agent Hallucinations: 5 Techniques + Production Patterns
The Factory That Dreams: 39 AI Agents, No Framework
Should AI Engineers Still Read Code in 2026? The Z/L Continuum
Understanding is the new bottleneck
The Golden Age of AI Engineering
Building an ACP-Compatible Agent Live
Everything We Knew About Software Has Changed
I Run a Fleet of AI Agents Across Three Machines. Here's What Broke.
Teaching Coding Agents to Do Spreadsheets
Think You Can Build a Game with AI? Think Again!
Your agent is blindfolded
Your coding agent doesn't always follow your rules
Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data
500 People Vibe-Coded for 30 Days. I Was One of Them.
Beyond the Harness: A Journey Towards Adaptive Engineering
Build AI Systems for Discernment, Not Approval
GTM Is You
How We Taught Agents to Use Good Retrieval
Respect The Process
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale
The Pipeline Is Dead
Field Guide to Fable
Continual Learning for AI Agents: From Failures to Durable Improvements
MCP Apps: Primitives, Discovery, and the Future of Software
The Missing Layer After Launch
Your AI Product Will Fail Unless You Can Explain It
The Prompt Is Still a Punch Card
Software Factories & Keynotes
Building Great Agent Skills: The Missing Manual
Deterministic Infra for Non-Deterministic AI Agents
Frontier Results, On Device
The Agentic AI Engineer
The Future Is Domain-Specific Agents
The Prompt is the Platform
You Can't Prompt the Room: The Last Skill AI Won't Replace
Your Agent Failed in Prod. Good Luck Reproducing It.
Agents Building Agents
AI-Driven Multi-Document Correlation for Financial Compliance
AI System Design: From Idea to Production
Browser Agents Don't Need Better Models. They Need Better Eyes.
Building an Autonomous Engineering Org
Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry
HTML is All You Need (for Agents to Make Graphics)
OpenClaw in Your Hand: Building a Physical AI Terminal
Research to Reality: Bringing Frontier ML Research to Production
Structuring the Unstructured
The 100-Tool Agent Is a Trap
User Signal Dies at the Retrieval Boundary
Using Spec-Driven Development for Production Workflows
Voice In, Visuals Out: The Agony and the Ecstasy
We Cut 94% of AI Coding Tokens With a Local Code Index
When All Context Matters: Extended Cache Augmented Generation
Your Agent Is Wasting Tokens and You Don't Know It
A Genius With Amnesia
Agents in Production: How OpenGov Built and Scaled OG Assist
Stop Writing Tone Instructions. Layer Them.
Turn 10,994 Notes Into Memory
Build Systems, Not Code
Production Evals For Agentic AI Systems
Recursive Coding Agents
The Log Is The Agent
The Miranda Hypothesis: How Hamilton Poisoned Persona Evals
6 Things to Know about AIE World's Fair 2026
The Production AI Playbook: Deploying Agents at Enterprise Scale
Your Agent's Biggest Lie: "I Searched the Web"
You Might Not Need 50 Diffusion Steps
Why MCP and ChatGPT Apps Use Double Iframes
The agent-ready web: Simplify user actions with WebMCP
Why Can't Anyone Answer Questions About the Business?
Your Attention Is the Bottleneck, Not Your Agents
Self Driving Products: Product Signals to Pull Requests
Sovereign Escape Velocity: Ownership with Open Models
Stop Making Models Bigger, Make Them Behave
2026 AI Engineer Vibe Reel
GPU Cloud Deployment Without Leaving Your IDE
RAG is dead, right??
Road to 5 Million Tokens: Breaking Barriers in Long Context Training
Why Eval++ Is the Next Great Compute Primitive
Why More Context Makes Your Agent Dumber and What to Do About It
From MCP to Scale: Pipelines That Build Themselves
LLM Observability, Evaluation, Experimentation Platform
Under 5 minutes to a deployed LLM endpoint
Building Interactive UIs in VS Code with MCP Apps
Building Safe Payment Infrastructure for the Autonomous Economy
Evals Are Broken, Use Them Anyway
Beyond Transcription: Building Voice AI That Understands Conversations
Building Agent Interfaces: Lessons from Chrome DevTools (MCP) for Agents
The Art & Science of Benchmarking Agents
Software Fundamentals Matter More Than Ever