evals
97 talks
Multimodal Collaborative Agents for Next-Gen Commerce
SOTA Generative Media Panel
Agents Are Where Microservices Were in 2015
The Half Life of Agent Infrastructure
AI Evals for Cross-Functional Teams
Building the Engine While Flying the Plane: Launching the Figma MCP Server
Building uReview, Uber's Multi-Agent Code Review Engine
How do you diffuse AI into the real world?
The Agentic Commerce Stack
Building GTM AI Agents: Lessons from Deploying to 6,000 Users
The Death of Developer Advocates
Preferences Over Benchmarks: Model Routing
What If Your Chip Design Team Moved Like a Single Body?
Prototyping as Leadership: How a CTO Ships with AI Agents
Your Agent Evolved. Your Evals Didn't.
200 Million Patient Interactions Later
AI is the World's Largest Relationship Therapist
Don't Be Data Poor
From Ambient Documentation to Clinical Intelligence
Guardrails First: Engineering Member-Facing Health AI
Shipping AI to a Million Patients Without an A/B Test
Trading Desks to Clinical Trials: Parallels in Applied Vertical AI
Why Your Enterprise Tech Stack Isn't Ready for AI Agents
Computer Use at the Edge of the Statistical Precipice
Beyond Static Intelligence: Evaluating Continual Learning
Designing Agents (The Floor Is the Frontier)
Improving Agents is a Data Mining Problem
Memory Harnesses for Long-Running Research Agents
Multiplayer agentic engineering
Benchmarking Coding Agents on New vs Legacy Codebases
Compression at the Edge
When Will the Benchmaxxing Plague End?
Rethinking Environments for Long-Horizon Work
Teaching AI to Find Real Vulnerabilities
Benchmarks: The Good, the Bad, and the Ugly
Ending AI Slop
Fighting Slop with Slop
Reinforcement Learning without Verifiable Rewards
Verifiable Environments for AI in Biology
Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains
Persona Engineering: A Field Guide to AI Synthetic Personas
SimulationMaxxing: How We Ship Agents 20× Faster
Skills are New Features: Building a Skill-Centric Harness
Why Off-the-Shelf AI Doesn't Understand Money
How Forward Deployed Engineering is done at Ramp
DeepSWE: A Contamination-Resistant Coding Benchmark
State of Data
Evaling Video Slop
Evals-Driven Development for a Mental Health AI Coach
From Agent Traces to Agent Simulations
Building Closed-Loop Evals for a Multimodal Agent at Scale
Everything Is a Rollout
How Evals and Prompts Shape Agent Behavior
The Future of Evals: From LLM as a Judge to Agent as a Judge
Training Frontier Models to Out-Think Hackers
Vending-Bench: Long-Horizon Agent Evals
The Unreasonable Effectiveness of Separating the Task from the Model
Active Graph Agent Runtime (BabyAGI 4)
Your Moat Is Your Data Model
The Biggest Challenge in Your Stack? Evals, Evals, Evals
Enterprise Agents Have a Structure Problem
Medic for Apache Spark: First Aid for Failing Jobs
Build Evals That Actually Matter
Stop Burning Tokens: Why Self-Improvement Needs Domain Expertise First
Your Agents Need a Save Button
On AI and Knowledge
Software engineering is not about writing code
Special Topics in Kernels, RL, Reward Hacking in Agents
Recursive Model Improvement
Don't Ship Skills Without Evals
Modern Post-Training: A Deep Dive
Stop Evaluating Models Like It's the 50s
Semantic Blindness: 500,000 Sensors Confused an LLM
Teaching Coding Agents to Do Spreadsheets
Your coding agent doesn't always follow your rules
Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data
Build AI Systems for Discernment, Not Approval
Respect The Process
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale
Continual Learning for AI Agents: From Failures to Durable Improvements
The Missing Layer After Launch
Software Factories & Keynotes
Frontier Results, On Device
The Agentic AI Engineer
Agents Building Agents
AI System Design: From Idea to Production
User Signal Dies at the Retrieval Boundary
Agents in Production: How OpenGov Built and Scaled OG Assist
Production Evals For Agentic AI Systems
The Miranda Hypothesis: How Hamilton Poisoned Persona Evals
The Production AI Playbook: Deploying Agents at Enterprise Scale
Why Can't Anyone Answer Questions About the Business?
Self Driving Products: Product Signals to Pull Requests
Stop Making Models Bigger, Make Them Behave
LLM Observability, Evaluation, Experimentation Platform
Evals Are Broken, Use Them Anyway
The Art & Science of Benchmarking Agents