# Long-Horizon Agents Need Experiments, Not Just Prompts

Erina Karati & Arunachalam Manikandan, Supercell | AI Engineer | 21:27

Source: https://www.youtube.com/watch?v=x4e5O9zN0TE
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/long-horizon-agents-need-experiments-not-just-prompts
Published: 2026-09-26
Tags: agents, evals, memory, multi-agent, planning

## TL;DR
- Long-horizon agents lose source attribution, uncertainty, and planning context even when they can retrieve memories.
- Controlled scenarios, trace collection, and a balanced scorecard let an autoresearch loop evaluate whole runs instead of single responses.
- The loop should edit a small policy surface and keep a change only when it improves the scorecard without breaking guardrails.

## Summary
Erina Karati presents Project Paradox, a stateful framework for game agents with per-agent memory, emotions, beliefs, trust scores, and actions in a shared world. The system handles short interactions well, but longer social chains expose failures: agents forget who originated information, turn rumors into facts, or fail to use remembered facts when planning. Karati wraps the village in an autoresearch loop. It runs controlled scenarios, records traces such as conversations and memory retrievals, scores diffusion, source retention, uncertainty, replanning, and privacy, then tests a constrained policy change. Changes are reverted when the scorecard or guardrails get worse. Karati is careful not to claim general improvement from the work. Her practical proposal is to freeze the harness, scenarios, and metrics, then expose only a small policy surface for search. She argues that the same pattern applies to support, research, coding, workflow, and personal assistant agents that maintain state over time.

## Key ideas
### Project Paradox gives game agents state that affects future actions
[00:12](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=12s)
Project Paradox is a modular framework for autonomous agents in a video game. Agents can move with intent, interact with objects and characters, react to events, hold conversations, and let those interactions affect their emotions, beliefs, and goals. Each agent has its own memory namespace backed by retrieval-augmented generation, so memories do not bleed between agents. The architecture also tracks emotion as a small vector and belief scores toward other agents and the player. A memory receives an importance score, and important events go into a separate cache for later retrieval.

### Short interactions work, but long social chains lose meaning
[05:08](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=308s)
Karati describes a mango-sale rumor that passes from one agent to another. After several intervening events, an agent may remember the general topic while losing who started it. The rumor can become a certain fact, or an agent can know a fact but fail to use it while planning an action. This is a failure of social consistency over time, rather than a failure to produce one plausible response. The system needs to preserve source, uncertainty, and the relationship between remembered information and later plans.

### Autoresearch evaluates complete runs instead of isolated answers
[06:43](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=403s)
Karati takes inspiration from Andrej Karpathy's autoresearch idea and applies it outside the village. Project Paradox becomes a lab bench, while the autoresearch layer reads full traces, compares behavior with scenario ground truth, proposes a constrained change to the agent protocol, and reruns the scenario. The layer is not another villager and does not give agents a shared memory. It evaluates what happened across the run and asks whether society-level behavior improved.

### Controlled scenarios make long-horizon behavior measurable
[10:43](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=643s)
Free-form village activity can look convincing while remaining difficult to evaluate. Karati proposes scenario suites with explicit tests. A public-fact scenario checks whether the right agents learn a fact, remember its source, and change their plans. A rumor scenario checks whether 'might leave' stays uncertain instead of becoming 'is leaving.' A replanning scenario checks whether agents update and communicate a blocked route. These scenarios give the loop repeatable conditions for comparing policy changes.

### A balanced scorecard prevents one metric from distorting behavior
[13:03](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=783s)
Karati argues against a single metric such as agent quality. The scorecard can measure reach for information diffusion, source retention for provenance, uncertainty preservation and false-send rate for rumors, action consistency and time to replan for planning, and containment for privacy. Optimizing diffusion alone could make agents overshare. Optimizing recall alone could create noisy or stale memories. Multiple measures keep the search from improving one score by damaging another.

### The editable policy surface must stay small
[14:33](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=873s)
The autoresearch layer should not rewrite the application or change the evaluation itself. Karati recommends freezing the harness, scenarios, and metrics, then exposing only policies such as memory writing, retrieval, communication, trust updates, source attribution, and replanning triggers. Example changes include preserving a source in memory, recording whether a claim is first- or secondhand, requiring hedging for uncertain claims, and proactively sharing important public facts. This limits direct gaming while leaving room to alter social behavior.

### Memory needs provenance and belief state, not only retrieval
[16:33](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=993s)
Karati is careful about the strength of her claims. She says the exposed surface is suitable for autoresearch because it is small enough to control and rich enough to affect social behavior, rather than claiming that the system generally improved. Her larger lesson is that adding retrieval-augmented memory does not solve long-horizon behavior. Agents may need to know whether information was firsthand, secondhand, verified, or uncertain, and they may need separate representations for raw episodic memories and current beliefs.

### Rollback protects against improvements that cause new failures
[17:48](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=1068s)
A policy can improve one behavior while damaging another. Faster public-fact diffusion might leak private information, while stronger recall might increase stale-memory use. Karati therefore describes the loop as a ratchet: try a change, score it, and retain it only if the scorecard improves and guardrails hold. Otherwise, revert to the earlier policy. This rollback rule is part of the experiment design, not an optional operational detail.

### The same state problem appears in many agent products
[18:23](https://www.youtube.com/watch?v=x4e5O9zN0TE&t=1103s)
Karati extends the pattern beyond games. Support agents need to track where a policy update came from and whether it supersedes an older answer. Personal assistants need to remember commitments and revise them when users change those commitments. Research agents need provenance, citation handling, contradiction handling, and hypothesis updates. Coding agents carry context across issues, files, teammates, and changing requirements. Workflow agents need access controls, handoffs, and replanning. Each system maintains state that affects later actions.

## Notable quotes
- Erina Karati: "The system may remember the rough topic but lose the source of the topic." (05:58)
- Erina Karati: "We were no longer evaluating one answer. We were evaluating an entire run." (09:04)
- Erina Karati: "The biggest lesson for me perhaps was that memory is not enough here." (16:56)
- Erina Karati: "Rollback also is not optional." (17:48)
- Erina Karati: "Long horizon agents need experiments and not just prompts." (20:41)

## Tools & references mentioned
- Project Paradox
- Supercell AI Innovation Lab
- Supercell
- autoresearch
- Andrej Karpathy
- Microsoft
- RAG

## Who should watch
- You are building agents that need to remember facts, commitments, or conversations across multiple steps and want tests beyond single-turn evaluations.
- Your agent sometimes loses provenance, hardens uncertain claims, or acts on stale state after the world changes.
- You need a practical structure for letting a model search over policy changes without giving it permission to rewrite the whole system.

## Editor's note

Erina Karati shows that agents can produce plausible responses while losing the source and certainty of information across a long social chain. Kitaru records each model call and tool result from a real run, then replays the agent against the same inputs and responses. That makes a failed run repeatable after a change to the prompt, model, or code, so it can become a test case rather than a one-off observation.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [Build Agents That Run for Hours](https://aietalks.com/talks/build-agents-that-run-for-hours) (Ash Prabaker & Andrew Wilson, Anthropic, 1:15:40)
- [How We Solved Context Management in Agents](https://aietalks.com/talks/how-we-solved-context-management-in-agents) (Sally-Ann DeLucia, Arize AI, 16:17)
- [Claude for Long-Horizon Tasks](https://aietalks.com/talks/claude-for-long-horizon-tasks) (Lance Martin, Anthropic, 25:19)
- [Designing Agents (The Floor Is the Frontier)](https://aietalks.com/talks/designing-agents-the-floor-is-the-frontier) (Ben Hylak, Raindrop, 19:46)
- [Rethinking Environments for Long-Horizon Work](https://aietalks.com/talks/rethinking-environments-for-long-horizon-work) (Rayan Garg, Theta Software, 21:15)
