debugging
9 talks
From Signal to PR: Anatomy of a Self-Improving Agent
The Future of Evals: From LLM as a Judge to Agent as a Judge
Learned Execution Graphs for Anomaly Detection & Drift in APIs
Medic for Apache Spark: First Aid for Failing Jobs
Your Agents Need a Save Button
Your agent is blindfolded
The Agentic AI Engineer
Your Agent Failed in Prod. Good Luck Reproducing It.
Evals Are Broken, Use Them Anyway