agents
202 talks
Beyond the Lethal Trifecta: Agentic Commerce on the Open Internet
Multimodal Collaborative Agents for Next-Gen Commerce
Teaching agents to pay
The End of the Static Screen: Architecting Intent-Driven UX
When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS
Why Your AI Agent Needs a Wallet: USDC and Nanopayments
Your Agent Just Authorized What?!
Agents Are Where Microservices Were in 2015
Tell the Robot What You Want
The Half Life of Agent Infrastructure
Tribal Dungeons of Global Shipping: AI Agents at Global Scale
AI-Native Organisations Run on Skills: How to Structure and Scale Them
Building uReview, Uber's Multi-Agent Code Review Engine
From AI-Assisted to AI-Native: Building a Frontier Development Team
How do you diffuse AI into the real world?
How to Avoid Disaster When Vibe-Coding a Billing Engine
How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage)
Your Code Has Bugs. Lean4 Has Proofs: Formal Verification for Engineers
How Anthropic Builds: Lessons from Labs
How to Generate Mergeable Code with a Context Engine
The Agentic Commerce Stack
AI in GTM at Notion
Building GTM AI Agents: Lessons from Deploying to 6,000 Users
GTM Engineering: The Technical Bits
How AI Agents Let GTM Teams Scale
How We Got LLMs to Recommend Our Open Source Library
Knowledge Systems: The New GTM Stack
The Building Blocks of GTM Orchestration
The Death of Developer Advocates
The Missing Layer in Agentic AI
Einstein Arena: Harnessing Collective Agent Intelligence for Open Science
Agent Frameworks Considered Harmful
FinOps for AI Agents: Who Spent All the Tokens?
Give the Agent a Budget, Not a Token
The Agent Behind the Curtain: Building the Oz Cloud Agent Platform
What If Your Chip Design Team Moved Like a Single Body?
Agentic SDLC at Uber
Building Agents Is Trivial Now, Context Is the Next Frontier
The Missing Layer: Design Taste in AI Agents
How I automate my own job at Hugging Face using agents
IT Admin for the AI Workforce
Prototyping as Leadership: How a CTO Ships with AI Agents
The Era of Compound Engineering
Your Agent Evolved. Your Evals Didn't.
Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards
200 Million Patient Interactions Later
Don't Be Data Poor
Healthcare's Agent Bytecode: X12 as the Harness for AI Agents
Building an Agentic Video Editor for Mass Consumer
The Next Game Engine Won't Have a Manual
How to Kill the Code Review
Security Firewall for Agents
Bringing agents onto the World Wide Web
Computer-use models will agentify the web, not APIs
The Dark Arts of Web Automation: Teaching Agents to Use Websites Like Humans
Bringing Continual Learning into Enterprises
Designing Agents (The Floor Is the Frontier)
Intelligence + Continual Learning = Expertise
LLM Knowledge Bases: A Practical Guide
Agents, codebases, and teams
The Evolution of Agentic Surfaces
Codex, Behind the Harness
Always-on agents run production without the on-call tax
Guide, Verify, Solve
Anthropic's CCA Exam as a Field-Guide for Agentic Engineering
Benchmarking Coding Agents on New vs Legacy Codebases
Realtime multiplayer, automation, and you!
The New Primitives: Building AI Native Software
The State of Model Routing
Rethinking Environments for Long-Horizon Work
Benchmarks: The Good, the Bad, and the Ugly
Fighting Slop with Slop
Learning on the Job: The Future of Post-Training
Reinforcement Learning without Verifiable Rewards
The Data for Fully Autonomous Software Engineers and Companies
Verifiable Environments for AI in Biology
First Steps Toward Automated AI Research
Let's integrate AI Agents in Event-Sourced Systems
Your Finance Agent's Bottleneck Is You
Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains
SimulationMaxxing: How We Ship Agents 20× Faster
Skills are New Features: Building a Skill-Centric Harness
Wearing the Agent: From Group Chats to Glasses
Your Agent Didn't Fail. Your Harness Did.
AI Agents for Performance: Ship Faster, Pay Less
AI tools for Forward Deployed Engineering
How Forward Deployed Engineering is Done at Cognition
How Forward Deployed Engineering Is Done at Decagon
How Forward Deployed Engineering is done at Factory
How Forward Deployed Engineering is done at Ramp
The Dirty Secret of Forward Deployed Engineering
DeepSWE: A Contamination-Resistant Coding Benchmark
Evaling Video Slop
Loop Engineering from First Principles
Building Closed-Loop Evals for a Multimodal Agent at Scale
Everything Is a Rollout
From Signal to PR: Anatomy of a Self-Improving Agent
Setting Yourself Up for Success
The Future of Evals: From LLM as a Judge to Agent as a Judge
Vending-Bench: Long-Horizon Agent Evals
AI on Your Lakehouse: Context Comes in Shapes, Not Queries
Harness Engineering is Not Enough: Why Software Factories Fail
Why Agentic Systems Need Ontologies
Why We Killed Our Multi-Agent Pipeline
Active Graph Agent Runtime (BabyAGI 4)
Claude for Long-Horizon Tasks
From Systems of Record to Systems of Context
Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer
Better Auth
Every Harness Will Become A Claw
HTML Is All Agents Need
The Biggest Challenge in Your Stack? Evals, Evals, Evals
Your agent architecture has a half-life of 6 months
Build the AI GTM Agent That Knows the Buyer
Can Oncology Workflows Run Without Human Touch?
Enterprise Agents Have a Structure Problem
In the Land of AI Agents, the Verifiers Are King
It's 10pm. Do You Know Where Your Agents Are?
Medic for Apache Spark: First Aid for Failing Jobs
Security Track Intro
Skills are the New SDKs
We Gave an Agent Production Code Access and Then Tried to Sleep at Night
When Agents Meet Physical Data: The Other Physics of Agent Harnesses
Why Your Agent Disagrees With Itself (And What To Do About It)
From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization
You Didn't Ship a Bug. You Just Wrote It for a Human.
Agents Need Feature Flags
Autonomous Agents for Scientific Tasks
Your Agents Need a Save Button
Every company should have a Brain
Special Topics in Kernels, RL, Reward Hacking in Agents
The Great Loops Debate
Using LLMs to Secure Source Code
Imagination Engineering: "Live in the future and then build what's missing."
Computer-Use 2.0: Agents Just Got Multi-Cursor
Recursive Model Improvement
Don't Ship Skills Without Evals
The engineer of the future is the person who is able to choose what is worth doing.
WTF Is the Context Layer? The Missing Infrastructure for Production Agents
I've Never Seen Anything Scarier Than an LLM with Tool Calls
A Song of Types and Agents
remobi.app: Don't Change Your Terminal Workflow for Mobile
ReviewDebt: A Practical Framework for Scoring Every Pull Request
Semantic Blindness: 500,000 Sensors Confused an LLM
The Agentic Web and the Bazaar Era of AI
The AI bugpocalypse is here. Now what?
What Does Done Even Mean? Agents and Paperclip's Liveness Model
Chat and citations won't save your vertical AI
Claws Out: Securing and Building with OpenClaw
Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers
Develop at Idea Velocity
Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD
From Writing Code to Designing Systems: How the Developer Role is Changing
Stop AI Agent Hallucinations: 5 Techniques + Production Patterns
The Factory That Dreams: 39 AI Agents, No Framework
Should AI Engineers Still Read Code in 2026? The Z/L Continuum
Understanding is the new bottleneck
The Golden Age of AI Engineering
Building an ACP-Compatible Agent Live
Everything We Knew About Software Has Changed
I Run a Fleet of AI Agents Across Three Machines. Here's What Broke.
Teaching Coding Agents to Do Spreadsheets
Think You Can Build a Game with AI? Think Again!
Your agent is blindfolded
Your coding agent doesn't always follow your rules
500 People Vibe-Coded for 30 Days. I Was One of Them.
Beyond the Harness: A Journey Towards Adaptive Engineering
Build AI Systems for Discernment, Not Approval
Respect The Process
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale
Field Guide to Fable
Continual Learning for AI Agents: From Failures to Durable Improvements
The Missing Layer After Launch
Software Factories & Keynotes
Building Great Agent Skills: The Missing Manual
Deterministic Infra for Non-Deterministic AI Agents
The Agentic AI Engineer
The Future Is Domain-Specific Agents
The Prompt is the Platform
You Can't Prompt the Room: The Last Skill AI Won't Replace
Your Agent Failed in Prod. Good Luck Reproducing It.
Agents Building Agents
Building an Autonomous Engineering Org
HTML is All You Need (for Agents to Make Graphics)
User Signal Dies at the Retrieval Boundary
Using Spec-Driven Development for Production Workflows
We Cut 94% of AI Coding Tokens With a Local Code Index
A Genius With Amnesia
Agents in Production: How OpenGov Built and Scaled OG Assist
Build Systems, Not Code
Recursive Coding Agents
The Production AI Playbook: Deploying Agents at Enterprise Scale
Your Attention Is the Bottleneck, Not Your Agents
Self Driving Products: Product Signals to Pull Requests
Why Eval++ Is the Next Great Compute Primitive
Why More Context Makes Your Agent Dumber and What to Do About It
Building Safe Payment Infrastructure for the Autonomous Economy
Evals Are Broken, Use Them Anyway
The Art & Science of Benchmarking Agents
Software Fundamentals Matter More Than Ever
Don't Build Agents, Build Skills Instead
No Vibes Allowed: Solving Hard Problems in Complex Codebases