benchmarks
16 talks
Can LLMs Write Fast Multi-GPU Kernels?
Einstein Arena: Harnessing Collective Agent Intelligence for Open Science
Computer Use at the Edge of the Statistical Precipice
Beyond Static Intelligence: Evaluating Continual Learning
When Will the Benchmaxxing Plague End?
Teaching AI to Find Real Vulnerabilities
Benchmarks: The Good, the Bad, and the Ugly
Scaling to Long Horizons
DeepSWE: A Contamination-Resistant Coding Benchmark
State of Data
From Agent Traces to Agent Simulations
Training Frontier Models to Out-Think Hackers
Computer-Use 2.0: Agents Just Got Multi-Cursor
Stop Evaluating Models Like It's the 50s
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale
The Art & Science of Benchmarking Agents