# Stop Rationing Tokens: Let the Harness Pick the Model

Laurent Gil & Žilvinas Urbonas, Cast AI | AI Engineer | 18:08

Source: https://www.youtube.com/watch?v=48YUYDjwfYY
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/stop-rationing-tokens-let-the-harness-pick-the-model
Published: 2026-10-02
Tags: agents, coding-agents, cost, evals, harness-engineering

## TL;DR
- Kimchi measures cost per completed task rather than cost per token, then selects a model that can deliver the required outcome at lower cost.
- After three months with 300 employees, Cast AI reports 2.5 times savings while token use increased by 1.5 times.
- Ferment runs coding tasks for hours, checks its own output, repeats work when the score is too low, and deploys completed changes to staging.

## Summary
Laurent Gil argues that companies should stop limiting developers' token use. Cast AI instead built Kimchi, an open-source coding harness that makes tokens cheaper by choosing a model according to the cost and quality of the completed task. The harness can change its choice as new models become available. Gil reports 2.5 times savings over three months with 300 employees, even though token use increased by 1.5 times. Žilvinas Urbonas explains how Kimchi runs longer tasks through milestones, scoring, rebuilds, and checks before sending changes to staging. The talk also introduces Teleport, which keeps agent sessions running in remote sandboxes after a laptop is closed, and Studio, which lets teams review and take over sessions through a shared board. The system still stops at staging because the speakers believe a person should review changes before production for now.

## Key ideas
### Token limits make coding agents harder to use
[01:16](https://www.youtube.com/watch?v=48YUYDjwfYY&t=76s)
Laurent Gil compares token limits to giving someone a laptop with one hour of battery and allowing a charge only once a day. He says managers should make tokens available for as long as developers need them instead of preventing developers from using coding agents. Cast AI built Kimchi around that goal. The internal user base included 300 employees, with about two-thirds working as developers, and the team had used the coding agent for three months when the talk was recorded.

### Cost per completed task gives a different picture from cost per token
[02:57](https://www.youtube.com/watch?v=48YUYDjwfYY&t=177s)
Gil shows a comparison of models that complete the same task at the same quality. He says Gemini 3 Flash appeared to cost $3.5 per million tokens in the study, while the task cost $75. He gives another example with MiniMax M2.7, whose token price was $1.5 while the task cost $148. His point is that a cheaper token does not necessarily produce a cheaper result. Kimchi therefore evaluates the cost of getting the desired task outcome.

### Kimchi changes model choice as model performance changes
[04:20](https://www.youtube.com/watch?v=48YUYDjwfYY&t=260s)
Kimchi's harness chooses a model for each task and can change that choice without a person retesting every new model. Gil shows a shift in the internal selection data. Kimi K2.6 was the most selected model on June 3, while MiniMax M3 was winning on June 21. He says an automated harness can track these changes because it evaluates the cost and outcome of the work instead of relying on a fixed model preference.

### The reported savings came while token use increased
[04:48](https://www.youtube.com/watch?v=48YUYDjwfYY&t=288s)
Gil reports 2.5 times savings over three months. The graph compares an estimated cloud bill with what Cast AI paid while achieving the same outcome. He says the number of tokens increased by 1.5 times, while the cost of coding decreased by 1.5 times. The harness made the coding agent effectively unlimited for the team because spending grew more slowly than token use for the tasks.

### Ferment breaks long coding tasks into scored milestones
[08:15](https://www.youtube.com/watch?v=48YUYDjwfYY&t=495s)
Žilvinas Urbonas describes Ferment as a construct for coding tasks that can run for more than two hours. It asks questions, breaks the work into milestones, implements the task, and scores the result. The score gives the team a way to judge whether the selected model produced an acceptable output. A result is considered complete when it reaches at least a B score, and the user can ask the agent to revise it toward an A score.

### Kimchi checks, rebuilds, and stops at staging
[09:01](https://www.youtube.com/watch?v=48YUYDjwfYY&t=541s)
The Kimchi workflow runs a build after a change and checks whether something broke. If it did, the system changes the code and rebuilds it again. Once the work is complete, it deploys the changes to a staging environment. Urbonas says the current design keeps a person in the loop before production. Future checks could include application health, Kubernetes pods, observability data, and other production metrics before an automatic production release.

### Teleport keeps agent sessions alive outside the laptop
[11:38](https://www.youtube.com/watch?v=48YUYDjwfYY&t=698s)
Kimchi Teleport moves an agent session into a secure remote sandbox so it can continue after the user closes a laptop. Gil explains that the local environment is synchronized with a container running in a Kubernetes cluster on Google infrastructure, so the session feels like the same environment while running elsewhere. Kimchi manages the remote session, including its size and security requirements, and users can access it from a laptop or mobile device when the agent needs clarification.

### Studio gives teams a shared board for agent work
[14:27](https://www.youtube.com/watch?v=48YUYDjwfYY&t=867s)
Kimchi Studio extends the remote-session idea to teams. It shows tasks, plans, work in progress, and sessions waiting for review in a browser-based board. An engineer can ask the harness for a plan, then a peer or product manager can review it before implementation continues. Team members can take over items in review, return them to progress, and start work from a backlog. Studio is built on Teleport sandboxes, while the harness is open source and the other tools require a Google account.

## Notable quotes
- Laurent Gil: "Our job is to make sure they can use it as much as they want for as long as they want in a completely unlimited fashion." (02:05)
- Laurent Gil: "Would it be nice that you have an automated engine, automated harness that will select the right model for the task at the right time based on the outcome of the task." (04:20)
- Žilvinas Urbonas: "Artifact is only considered complete once the output of the scoring is at least B as a score." (10:06)
- Žilvinas Urbonas: "Kimchi Teleport allows you to close your laptop." (11:38)
- Laurent Gil: "Studio is a very nice way to visualize all the sessions that are running including those that are asking for a review." (15:40)

## Tools & references mentioned
- Cast AI
- Kimchi
- Kimchi Coding
- Ferment
- Kimchi Teleport
- Kimchi Studio
- Claude Code
- Anthropic
- Gemini 3 Flash
- MiniMax M2.7
- MiniMax M3
- Kimi K2.6
- Google
- Kubernetes

## Who should watch
- You are paying more for coding agents as developers use more tokens, and you want to reduce spend without imposing token limits.
- You are building an internal coding harness and need model selection to follow task outcomes rather than a permanent model choice.
- Your coding agents need to run for hours, survive laptop shutdowns, or move through review with other engineers and product staff.

## Related talks

- [FinOps for AI Agents: Who Spent All the Tokens?](https://aietalks.com/talks/finops-for-ai-agents-who-spent-all-the-tokens) (Tisha Chawla & Susheem Koul, Microsoft, 21:24)
- [Tokens Should Have Jobs](https://aietalks.com/talks/tokens-should-have-jobs) (Katelyn Lesse & Angela Jiang, Anthropic, 13:21)
- [Your Agent Is Wasting Tokens and You Don't Know It](https://aietalks.com/talks/your-agent-is-wasting-tokens-and-you-dont-know-it) (Erik Hanchett, Amazon Web Services, 05:55)
- [Give the Agent a Budget, Not a Token](https://aietalks.com/talks/give-the-agent-a-budget-not-a-token) (Sachin Malhotra, Anthropic, 19:53)
- [Build Agents That Run for Hours](https://aietalks.com/talks/build-agents-that-run-for-hours) (Ash Prabaker & Andrew Wilson, Anthropic, 1:15:40)
