Evals provide ahead-of-time reliability, while harnesses protect an AI system at runtime.
2
An LLM judge needs human-scored examples, policy context, and calibration before it can gate a release.
3
A green evaluation can still be wrong because judges favor position, agreeable answers, their own model family, and verbosity.
Summary
Tejas Kumar explains evaluations as reliability instrumentation for non-deterministic AI systems. He compares them with unit tests, then builds a customer-support example where a substring check passes an answer because it contains "cannot", even when the answer is wrong. The example becomes an LLM judge backed by scenario data, human verdicts, and a return policy. Kumar shows how the judge can prefer its own model's answer and how missing policy context creates more failures. He improves the result by examining disagreements, adding policy details, and changing the model when the cheaper model reaches its limits. He recommends running deterministic checks first, then trace and pattern checks, and using model judgments only when needed. A CI gate can ship when judge and human agreement reaches 80%, while fresh production traffic, human review, synthetic cases, and red-team prompts keep the dataset current. OpenRAG retrieves a living refund policy for both agents and judges.
Evals provide ahead-of-time reliability alongside runtime harness protection
Kumar distinguishes the reliability provided by a harness from the reliability provided by evals. A harness runs when a user makes a request, such as catching prompt injection in a web page returned by a tool. An eval runs before deployment and checks whether the system behaves acceptably across many scenarios. He compares this with programming languages: just-in-time behavior can fail at runtime, while ahead-of-time checks can stop a bad build. He also says evals can expose unsafe agent behavior before it reaches users, especially when an agent has access to account mutations or other sensitive actions.
An eval is a fuzzy unit test for a non-deterministic system
Kumar defines an eval as reliability instrumentation for systems whose outputs are not deterministic. A unit test can assert that add(1, 2) returns 3, but an AI application needs a less exact question: did the agent behave acceptably across a scenario? A production eval setup has test cases, expected behavior, a judge, and an aggregation method. The expected behavior is a direction rather than an exact string. The judge decides whether the response followed policy, used the right tools, or resolved the issue. Tool calls stored as JSON can still be checked with ordinary deterministic assertions.
Deterministic checks should come before model-based judgments
Kumar lays out several evaluation techniques. Exact matching can check whether a specific tool was called with particular arguments, and schema validation can inspect JSON or message envelopes. Pair-wise comparison asks a judge to choose between two answers, but Kumar prefers point-wise checks when possible because they ask whether one answer is acceptable. He cites a study in which pair-wise comparison was more vulnerable to attack vectors than point-wise comparison. An LLM judge can handle fuzzy quality judgments, but it is itself non-deterministic and needs to be tested. Later, Kumar recommends starting with unit tests, pattern matching, and trace inspection before paying for model judgments.
LLM judges can produce a green result for the wrong reasons
Kumar names four judge biases. Position bias makes a judge favor the first option, which can be tested by swapping answer order. Sycophancy makes a judge prefer the answer that sounds nicest, such as approving a return outside the stated policy. Self-preference makes a judge favor text produced by its own model family, so the judge should be from a different family when possible. Verbosity can make a longer wrong answer look better because the judge prefers more tokens. These failures mean a green eval is not proof that the system works. The judge itself needs calibration and adversarial checks.
Kumar builds a dataset containing a customer scenario, an agent answer, and a human verdict. He recommends starting with around 30 examples, then asking the judge to score the scenarios without showing it the human verdict. The team compares the judge's results with the human labels. Kumar sets a useful agreement range around 80 to 85 percent. A judge that agrees 100 percent may be overfit or sycophantic, while agreement below 80 percent is not reliable enough for the intended use. Disagreements become evidence about missing context, unclear policy, or model quality.
A CI gate turns agreement into a release condition
Kumar recommends running the calibrated evals in CI and blocking a release when agreement falls below 80 percent. The judge model should be pinned to a specific version instead of relying on an alias that may change underneath the test. Passing CI is only the start. Production traffic should be sampled, and an internal team should score new inputs and outputs so the human-labelled dataset remains current. Kumar also says the gate should have an upper bound: 100 percent agreement can indicate overfitting. The production harness still protects users at runtime when the pre-deployment checks miss something.
Policy retrieval fixes missing context before changing models
Kumar says policies change and should be retrieved from a living source rather than copied into prompts by hand. In the live example, the judge initially knows only that returns are allowed within 14 days. It then misses details such as used items, original packaging, and exceptions for loyal customers. Kumar adds those details as context and checks each disagreement. When the lower-quality model still fails after receiving the relevant policy, he changes the model. He demonstrates OpenRAG retrieving a newly uploaded refund policy, allowing the judge to deny a headphone return made after 20 days.
Evaluation datasets need synthetic and adversarial examples
In the questions, Kumar explains how to expand an eval dataset without waiting for customer traffic. Agents can generate synthetic customer scenarios from policy and existing data. Teams can ask them to break the system by creating cases that are likely to fail. Red teams use the same approach with adversarial prompts, and failed cases can be used to generate more cases with similar risks. Kumar also says the production agent and the judge should use the same living policy source. Sharing policy through retrieval gives both systems the same information while preserving the judge's similarity to the production agent.