reliability
37 talks
AI Agents Are Just Distributed Systems Now
From Tokenmaxxing to Trusted Throughput
Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons
How Anthropic Builds: Lessons from Labs
Give the Agent a Budget, Not a Token
Your Agent Evolved. Your Evals Didn't.
Infra behind Krea 2: How to train and serve at scale
Computer Use at the Edge of the Statistical Precipice
The Evolution of Agentic Surfaces
Benchmarking Coding Agents on New vs Legacy Codebases
MCP Tasks (async): Why Aren't Any Agents Supporting Them?
Data and Environment Curation for Post-Training LLMs
What's Next After RLHF?
Your Agent Didn't Fail. Your Harness Did.
The Messy Reality of Scale: Synthetic Data and Pre-Training
Learned Execution Graphs for Anomaly Detection & Drift in APIs
Perception Agents
Active Graph Agent Runtime (BabyAGI 4)
Your agent architecture has a half-life of 6 months
Build the AI GTM Agent That Knows the Buyer
Can Oncology Workflows Run Without Human Touch?
Don't Let the LLM Drive
In the Land of AI Agents, the Verifiers Are King
We Gave an Agent Production Code Access and Then Tried to Sleep at Night
Why Your Agent Disagrees With Itself (And What To Do About It)
The Great Loops Debate
From fork() to Fleet: Designing an Agent Sandbox Cloud
Stop Evaluating Models Like It's the 50s
What Does Done Even Mean? Agents and Paperclip's Liveness Model
Develop at Idea Velocity
Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD
Should AI Engineers Still Read Code in 2026? The Z/L Continuum
Deterministic Infra for Non-Deterministic AI Agents
The Prompt is the Platform
Production Evals For Agentic AI Systems
Recursive Coding Agents
The Log Is The Agent