
Why first: Khial shows how a benchmark lies in ordinary engineering terms. A test can demand an unstated variable name, a prompt can reveal the implementation, and an agent can search for the answer. Vidal follows by questioning the number produced even when every item has been graded correctly.







