testing
39 talks
From AI-Assisted to AI-Native: Building a Frontier Development Team
How to Avoid Disaster When Vibe-Coding a Billing Engine
How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage)
Your Code Has Bugs. Lean4 Has Proofs: Formal Verification for Engineers
AI is the World's Largest Relationship Therapist
How to build an AI-Native Health Company
Shipping AI to a Million Patients Without an A/B Test
How to Kill the Code Review
Designing Agents (The Floor Is the Frontier)
Guide, Verify, Solve
Building Turbopuffer
When Will the Benchmaxxing Plague End?
Benchmarks: The Good, the Bad, and the Ugly
SimulationMaxxing: How We Ship Agents 20× Faster
We Vetted 2,000 AI Skills Before They Reached Developers
AI Agents for Performance: Ship Faster, Pay Less
DeepSWE: A Contamination-Resistant Coding Benchmark
Evals-Driven Development for a Mental Health AI Coach
From Agent Traces to Agent Simulations
How Evals and Prompts Shape Agent Behavior
In the Land of AI Agents, the Verifiers Are King
Build Evals That Actually Matter
Stop Burning Tokens: Why Self-Improvement Needs Domain Expertise First
The Great Loops Debate
Don't Ship Skills Without Evals
Stop Evaluating Models Like It's the 50s
Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD
Should AI Engineers Still Read Code in 2026? The Z/L Continuum
Your agent is blindfolded
Your coding agent doesn't always follow your rules
Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data
The Pipeline Is Dead
Continual Learning for AI Agents: From Failures to Durable Improvements
The Prompt is the Platform
Your Agent Failed in Prod. Good Luck Reproducing It.
Your Attention Is the Bottleneck, Not Your Agents
LLM Observability, Evaluation, Experimentation Platform
Evals Are Broken, Use Them Anyway
Software Fundamentals Matter More Than Ever