AI Engineer World's Fair 2025
297 talks
Shipping AI That Works: An Evaluation Framework for PMs
From Vibe Coding To Vibe Engineering
Government Agents: AI Agents Meet Tough Regulations
VoiceVision RAG: Integrating Visual Document Intelligence with Voice Response
Challenges in High Performance Robotics Systems
Building an Agentic Platform
Five hard earned lessons about Evals
How BlackRock Builds Custom Knowledge Apps at Scale
Perceptual Evaluations: Evals for Aesthetics
Form factors for your new AI coworkers
Fuzzing in the GenAI Era
Multi Agent AI and Network Knowledge Graphs for Change
Wisdom-Driven Knowledge Augmented Generation at Scale
The Next Unicorns: 7 Top AI Startups from the HF0 Residency
#define AI Engineer
Designing AI-Intensive Applications
The Future of Evals
2025 is the Year of Evals! Just like 2024, and 2023, and …
Evals Are Not Unit Tests
How to Look at Your Data
On Engineering AI Systems that Endure the Bitter Lesson
Vibe Coding with Confidence
How to Improve Your Vibe Coding
Practical Tactics to Build Reliable AI Apps
Real World Development with GitHub Copilot and VS Code
Realtime Voice AI
Vibes won't cut it
Vision AI in 2025
Building Agents at Cloud Scale
State of Startups and AI 2025
Useful General Intelligence
Agents vs Workflows: Why Not Both?
Hacking the Inference Pareto Frontier
Infrastructure for the Singularity
The 2025 AI Engineering Report
Why We Don't Need More Data Centers
Building Conversational AI Agents
From Self-driving to Autonomous Voice Agents
Pipecat Cloud: Enterprise Voice Agents Built On Open Source
Serving Voice AI at $1/hr: Open-Source, LoRAs, Latency, Load Balancing
Why ChatGPT Keeps Interrupting You
Your realtime AI is ngmi
How to Defend Your Sites from AI Bots
How to Secure Agents using OAuth
How We Hacked YC Spring 2025 Batch's AI Agents
Securing Code-Executing AI Agents
The Unofficial Guide to Apple's Private Cloud Compute
Building a Smarter AI Agent with Neural RAG
Building Alice's Brain: An AI Sales Rep that Learns Like a Human
Building Metrics that Actually Work
Evaluating AI Search: A Practical Framework for Augmented AI Systems
Layering Every Technique in RAG, One Query at a Time
Scaling Enterprise-Grade RAG: Lessons from Legal Frontier
Building the Platform for Agent Coordination
Everything is ugly, so go build something that isn't
Make Your LLM App a Domain Expert: How to Build an Expert System
Real-time Experiments with an AI Co-Scientist
Scaling AI Agents Without Breaking Reliability
Shipping Products When You Don't Know What They Can Do
Shipping something to someone always wins
What Is a Humanoid Foundation Model? An Introduction to GR00T N1
Why Your Product Needs an AI Product Manager, and Why It Should Be You
Information Retrieval from the Ground Up
Ship Agents that Ship: A Hands-On Workshop
Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop)
The AI Engineer's Guide to Raising VC
Why you should care about AI interpretability
A2A & MCP Workshop: Automating Business Processes with LLMs
Introduction to LLM serving with SGLang
Piloting agents in GitHub Copilot
Robotics: why now?
Waymo's EMMA: Teaching Cars to Think
Beyond the Prototype: Using AI to Write High-Quality Code
Devin 2.0 and the Future of SWE
Human seeded Evals
Latent Space Paper Club: AIEWF Special Edition (Test of Time, DeepSeek R1/V3)
Ship Production Software in Minutes, Not Months
Software Development Agents: What Works and What Doesn't
Your Coding Agent Just Got Cloned And Your Brain Isn't Ready
AI That Pays: Lessons from Revenue Cycle
AX Is the Only Experience That Matters
Building AI Products That Actually Work
Building Applications with AI Agents
How to Build Enterprise-Aware Agents
Mentoring the Machine
Rise of the AI Architect
Structuring a Modern AI Team
The Rise of Open Models in the Enterprise
3 ingredients for building reliable enterprise agents
Build Dynamic Products, and Stop the AI Sideshow
Building Agents (the hard parts!)
Does AI Actually Boost Developer Productivity? (100k Devs Study)
From Copilot to Colleague: Trustworthy Agents for High-Stakes
From Hype to Habit: How We're Building an AI-First SaaS Company
How agents will unlock the $500B promise of AI
How Intuit Uses LLMs to Explain Taxes to Millions of Taxpayers
How to Build Planning Agents without Losing Control
Machines of Buying and Selling Grace
Monetizing AI
POC to PROD: Hard Lessons from 200+ Enterprise GenAI Deployments
The Billable Hour is Dead; Long Live the Billable Hour
AI-powered entomology: Lessons from millions of AI code reviews
Books reimagined: AI to create new experiences for things you know
Continuous Profiling for GPUs
Critical AI Inference Your CIO Can Trust
How to Hire AI Engineers when EVERYONE is cheating with AI
How to Run Evals at Scale: Thinking Beyond Accuracy or Similarity
HybridRAG: A Fusion of Graph and Vector Retrieval
Knowledge Graphs in Litigation Agents
Practical GraphRAG: Making LLMs Smarter with Knowledge Graphs
Stateful environments for vertical agents
Stop Using RAG as Memory
Top Ten Challenges to Reach AGI
When Vectors Break Down: Graph-Based RAG for Dense Enterprise Knowledge
AI and Human Whiteboarding Partnership
CIAM for AI: Authn/Authz for Agents
Good design hasn't changed with AI
The Bitter Layout, or How I Learned to Love the Model Picker
tldraw.computer
Building Effective Voice Agents
Robots as Professional Chefs
What Every AI Engineer Needs to Know About GPUs
A Taxonomy for Next-gen Reasoning
ComfyUI Full Workshop
Design like Karpathy is watching
Dream Machine: Scaling to 1m users in 4 days
Google Photos Magic Editor: GenAI Under the Hood of a Billion-User App
How to Train Your Agent: Building Reliable Agents with RL
On Curiosity
OpenThoughts: Data Recipes for Reasoning Models
Real-world MCPs in GitHub Copilot Agent Mode
Reinforcement Learning, Kernels, Reasoning, Quantization & Agents
Full Spec MCP: Hidden Capabilities of the MCP Spec
Shipping an Enterprise Voice AI Agent in 100 Days
The Rise of the Agentic Economy on the Shoulders of MCP
360Brew: LLM-based Personalized Ranking and Recommendation
Improving Recommendation Systems and Search in the Age of LLMs
Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3
One Model to Rule Netflix Recommendations
RL for Autonomous Coding
Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+DAU
The State of Generative Media
Transforming search and discovery using LLMs
What We Learned from Using LLMs in Pinterest
Benchmarks Are Memes: How What We Measure Shapes AI and Us
Bolt.new: How We Scaled $0-20M ARR in 60 Days with 15 People
Building a 10 person unicorn
Rethinking Team Building: How a 30-Person Startup Serves 50 Million Users
Small AI Teams with Huge Impact
Automating Escrow with USDC and AI
Prompt Engineering and AI Red Teaming
How LLMs Work for Web Devs: GPT in 600 Lines of Vanilla JS
AI Pipelines and Agents in Pure TypeScript with Mastra.ai
AI Engineering with the Google Gemini 2.5 Model Family
The New Code
A year of Gemini progress + what comes next
Production software keeps breaking and it will only get worse
Thinking Deeper in Gemini
2025 in LLMs so far, illustrated by Pelicans on Bicycles
Trends Across the AI Frontier
Training Agentic Reasoners
New York Times' Connections: A Case Study on NLP in Word Games
Claude Code & the evolution of agentic coding
12-Factor Agents: Patterns of Reliable LLM Applications
MCP Is Not Good Yet
The Build-Operate Divide: Bridging Product Vision and AI Operational Reality
Your Personal Open-Source Humanoid Robot for $8,999
Conquering Agent Chaos
Mastering AI Evaluation: From Playground to Production
Optimizing Inference for Voice Models in Production
The New Lean Startup
Agents, Access, and the Future of Machine Identity
Intro to GraphRAG
Securing Agents with Open Standards
The Emerging Skillset of Wielding Coding Agents
Turning Fails into Features: Zapier's Hard-Won Eval Lessons
Building voice agents with OpenAI
Containing Agent Chaos
Agentic Excellence: Mastering AI Agent Evals with Azure AI Evaluation SDK
Agentic GraphRAG: Simplifying Retrieval Across Structured and Unstructured Data
AI Red Teaming Agent: Azure AI Foundry
Architecting Agent Memory: Principles, Patterns, and Best Practices
Building agent fleet architectures your CISO doesn't hate
Building Agentic Applications with Heroku Managed Inference and Agents
Building Code First AI Agents with Azure AI Agent Service
Building Multimodal AI Agents From Scratch
CI in the Era of AI: From Unit Tests to Stochastic Evals
Collaborating with Agents in Your Software Development Workflow
Data is Your Differentiator: Building Secure and Tailored AI Systems
"Data readiness" is a Myth: Reliable AI with an Agentic Semantic Layer
Don't get one-shotted: Use AI to test, review, merge, and deploy code
Effective agent design patterns in production
Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work
Evals 101
Events are the Wrong Abstraction for Your AI Agents
Forget RAG Pipelines: Build Production-Ready Agents in 15 Minutes
Foundry Local: Cutting-Edge AI Experiences on Device with ONNX Runtime/Olive
From Mixture of Experts to Mixture of Agents with Super Fast Inference
Graph Intelligence: Enhance Reasoning and Retrieval Using Graph Analytics
GraphRAG methods to create optimized LLM context windows for Retrieval
How fast are LLM inference engines anyway?
How to build world-class AI products
Introducing Strands Agents, an Open Source AI Agents SDK
Mastering Engineering Flow with Windsurf
Memory Masterclass: Make Your AI Agents Remember What They Do!
Milliseconds to Magic: Real-Time Workflows Using the Gemini Live API and Pipecat
Prompt Engineering is Dead
RAG in 2025: State of the Art and the Road Forward
Realtime Conversational Video with Pipecat and Tavus
Revenue Engineering: How to Price (and Reprice) Your AI Product
Serving Voice AI at Scale
Ship it! Building Production Ready Agents
Taming Rogue AI Agents with Observability-Driven Evaluation
The Agent Awakens: Collaborative Development with Copilot
The Eyes Are The (Context) Window to The Soul: How Windsurf Gets to Know You
The State of AI-Powered Search and Retrieval
To the moon! Navigating deep context in legacy code with Augment Agent
Unlocking AI-Powered DevOps Within Your Organization
Vector Search Benchmark[eting]
Vibe Coding at Scale: Customizing AI Assistants for Enterprise Environments
Vibe Coding at Scale: Customizing AI Assistants for Enterprise Environments
What Does Enterprise-Ready MCP Mean?
Why should anyone care about Evals?
Why Your Agent's Brain Needs a Playbook: Practical Wins from Using Ontologies
Fun Stories from Building OpenRouter and Where All This Is Going
Building AI Agents that Actually Automate Knowledge Work
RFT, DPO, SFT: Fine-tuning with OpenAI
Windsurf everywhere, doing everything, all at once
Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP
Building Agents with Amazon Nova Act and MCP
Veo 3 for Developers
Building Protected MCP Servers
The State of MCP Observability: Observable.tools
The Web Browser Is All You Need
Remote MCPs: What We Learned from Shipping
The Geopolitics of AI Infrastructure
MCP: Origins and Requests For Startups
How to Build Trustworthy AI
Exposing Agents as MCP Servers with mcp-agent
Beyond Conversation: Why Documents Transform Natural Language into Code
Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents
Just do it. (let your tools think for themselves)
MCPs are Boring (or: Why We Are Losing the Sparkle of LLMs)
Supercharging developer workflow with Amazon Q Developer
The Many Ends of Programming
Why Bolt.new Won and Most DevTools AI Pivots Failed
AI Engineer World's Fair 2025, Day 2 Keynotes & SWE Agents Track
Evals
Reasoning + RL
AI Engineer World's Fair 2025, Day 1 Keynotes and MCP Track
Retrieval + Search
Tiny Teams
GraphRAG
LLM Recommendation Systems (RecSys)
The 4 Patterns of AI Native Development
7 Habits of Highly Effective Generative AI Evaluations
Agentic Enterprise: What Your CEO Must Know About AI
Agents reported thousands of bugs, how many were real?
Analyzing 10,000 Sales Calls With AI In 2 Weeks
Are MCPs Overhyped? A Rant about MCPs
Arrakis: How to Build an AI Sandbox from Scratch
Blender MCP and The Future Of Creative Tools
Breaking the Chain: Agent Continuations for Resumable AI Workflows
Building Reliable Support Agents Using the Effect TypeScript Library
Buy Now, Maybe Pay Later: Dealing with Prompt-Tax While Staying at the Frontier
ChatGPT is poorly designed. So I fixed it
Designing AI To Scale Human Thought
Effective AI Agents Need Data Flywheels, Not The Next Biggest LLM
From PM at Stripe to Building an AI Startup: A Recent Founder's Journey
GPU-less, Trust-less, Limit-less: Reimagining the Confidential AI Cloud
Grounded Reasoning Systems for Cloud Architecture
How agents broke app-level infrastructure
Invisible Users, Invisible Interfaces: Accelerating Design Iteration with AI Simulation
Letting AI Interface with Your App with MCP
Luminal: Search-Based Deep Learning Compilers
MCP Agent Fine-tuning Workshop
My AI Thinks I'm Eating My Feelings (and Other Nutritional Insights)
open-rag-eval: RAG Evaluation without "golden" answers
RAG Evaluation Is Broken! Here's Why (And How to Fix It)
Real AI Agents Need Planning, Not Just Prompting
Rust is the language of the AGI
Stop Ordering AI Takeout: A Cookbook for Winning When You Build In House
Text-to-Speech Data Preparation and Fine-tuning Workshop
The Agent Native Company
The Benchmarks Game: Why It's Rigged and How You Can (Really) Win
The Coherence Trap: Why LLMs Feel Smart (But Aren't Thinking)
The Current State of Browser Agents
The Demo I Wish I'd Had: OpenAI's Agents SDK... serverless!
The End of Awkward AI Transcriptions
The Future of Qwen: A Generalist Agent Model
The Knowledge Graph Mullet: Trimming GraphRAG Complexity
The RAG Stack We Landed On After 37 Fails
The Robots Are Coming for Your Job, and That's Okay
The Voice-First AI Overlay: Designing Conversational Co-Pilots
Unlocking Africa's Potential with AI
Why the Best AI Agents Are Built Without Frameworks
Will Agent Evaluation via MCP Stabilize Agent Networks?