# Every AI Company Is Accidentally Building a Bank

Dor Sasson, Stigg | AI Engineer World's Fair 2026 | 20:08

Source: https://www.youtube.com/watch?v=cf2IhzqeQH4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/every-ai-company-is-accidentally-building-a-bank
Published: 2026-10-07
Tags: cost, enterprise, finance

## TL;DR
- AI pricing emergencies often come from infrastructure that checks usage after consumption instead of before it.
- AI products need synchronous entitlement checks before inference, followed by asynchronous settlement and reconciliation.
- Credit pools, concurrent agents, budgets, and enterprise controls require financial-system patterns such as holds, double-entry bookkeeping, and spend caps.

## Summary
Dor Sasson argues that AI products are becoming financial systems because their costs depend on usage that can change at runtime. He points to a series of April pricing emergencies involving Anthropic, OpenAI, and GitHub, along with companies that burned through large AI credit pools in weeks. In his view, these incidents happened because products allowed consumption and reconciled entitlements only after the invoice. Sasson proposes checking access and available balances synchronously before inference, then settling actual usage asynchronously. He compares this to an ATM approving a withdrawal before dispensing cash. He also covers concurrency, double spending, credit sources, hierarchical budgets, and spend caps across organizations, teams, users, and agents. The talk presents four recurring patterns for AI infrastructure, including reserve-and-settle, in-VPC metering, agent balance checks, and consumption visibility.

## Key ideas
### April's pricing emergencies exposed infrastructure failures
[01:17](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=77s)
Sasson describes April as a month when several code-generating products struggled with a new usage economy. Anthropic removed access for OpenClaw and other third-party agents from subscription plans. OpenAI changed product pricing overnight, and GitHub froze and then removed free Copilot trial access. He says these incidents were more than commercial problems. They showed what happens when products scale without the underlying controls needed for unpredictable AI consumption. He also cites companies burning through large AI cloud-credit pools, Uber using its yearly AI budget within weeks, and Replit illustrating how three users could consume an organization's entire credit pool.

### Anthropic could not separate subscription use from API consumption
[02:52](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=172s)
Sasson uses Anthropic and OpenClaw as a specific example. Users paid subscription-level dollars and cents, while Anthropic could incur much higher costs for the corresponding consumption. He says Anthropic lacked an effective way to distinguish people using its APIs from people using subscription programs. That left the company with an immediate need to stop access. The visible problem looked like pricing, but Sasson describes the underlying issue as infrastructure. The product did not have the runtime controls needed to prevent one type of entitlement from subsidizing another type of workload.

### Entitlements need to be checked before inference
[05:02](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=302s)
Sasson says the common failure in these incidents was timing. Users and agents could consume the product and burn tokens, while the entitlement check happened only after the invoice. By then, the spend and its consequences were already fixed. He proposes a different architecture: evaluate access synchronously on the request path, before inference and before balances are updated, then reconcile and settle the actual usage asynchronously afterward. The synchronous decision can include the user's plan, feature access, rate limits, trial status, and promotional rules. Sasson says this decision must happen before consumption because an invoice cannot prevent overspend.

### An ATM illustrates the required control point
[07:32](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=452s)
Sasson compares AI consumption to withdrawing cash from an ATM. The bank checks whether the account permits the withdrawal before handing over the money. It does not dispense cash and decide afterward whether the customer was entitled to it. He says AI products are missing this financial infrastructure layer. OpenAI's financial engineering team, he explains, has published an architecture with a runtime decision waterfall. That system evaluates the client, user, or agent against multiple policies and determines whether it can access a feature, operate under a rate limit, or draw down from a balance. Sasson describes the evaluation as synchronous and deterministic.

### Concurrent agents create a double-spend problem
[10:31](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=631s)
A shared balance can appear sufficient when many requests inspect it at the same time. Sasson imagines thousands of agents trying to draw from one dollar pool simultaneously. Without a hold-and-settle mechanism, each request may see available balance before the others have claimed it. He connects this to the classic double-spend problem and to double-entry bookkeeping. AI systems need ways to reserve funds, settle actual usage, and make requests idempotent and auditable. He argues that these concerns become more important as revenue and usage grow, and that they are difficult to add after a product has already been built around simpler consumption records.

### Credits need sources and explicit drawdown rules
[12:16](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=736s)
Sasson says a credit balance is rarely a single interchangeable number. A company may have credits from a debit account, cash account, savings account, promotion, grant, or another source. Each pool can have different rules. When a user consumes credits, the system needs to decide which pool is used first, such as older credits or promotional credits. He argues that many companies currently try to represent this with a single integer in a database, which does not capture the meaning of each source. The drawdown decision becomes part of the financial logic of the product.

### Enterprise AI budgets need hierarchy and spend caps
[13:31](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=811s)
Sasson says software consumption used to be relatively flat, with users and organizations as the main dimensions. AI introduces more layers. A contract may be approved at the executive level, while usage is distributed across teams, users, agents, models, features, and products. Buyers want to know who used the credits and to control how much each group can spend. That requires allocations, budgets, governance rules, and spend caps at different levels. He also describes the need to examine consumption across many dimensions, so finance and technology leaders can see how workloads are using the contract rather than receiving only a total usage number.

### Four infrastructure patterns recur in AI products
[14:46](https://www.youtube.com/watch?v=cf2IhzqeQH4&t=886s)
Sasson groups the emerging architecture into four recurring patterns. The first is reserving credit before inference and settling actual usage afterward. The second is keeping usage events, the ledger, and metering inside a company's virtual private cloud because of latency, cost, or data-sovereignty concerns. The third is checking balances as agents continue working, since their eventual task cost is unknown in advance, while reserving an account and reconciling later. The fourth is consumption visibility across models, features, and products. Sasson says this visibility is becoming an expected part of selling AI workloads, especially for finance and technology leaders evaluating enterprise usage.

## Notable quotes
- "Every AI company right now is accidentally building a bank." (00:31)
- "These pricing emergencies go way beyond just emergencies from a financial commercial set of the house. These really are infrastructure emergencies." (01:42)
- "The checks for what you're entitled, what you're allowed to do, whether you're an agent or a user, were happening after the invoice and not before." (04:58)
- "When you draw the cash, the check if you are allowed to draw it doesn't happen after you get the cash. It happens before you draw the cash." (07:32)
- "I don't think this is billing. I don't think this is necessarily this entire architecture is just related to the idea of invoicing and billing. This is a financial system that actually makes decisions ahead of any invoice." (09:41)

## Tools & references mentioned
- Stigg
- Anthropic
- OpenClaw
- OpenAI
- GitHub Copilot
- Uber
- Replit

## Who should watch
- You are building an AI product with subscriptions, usage pricing, or third-party agents and need to stop overspend before inference runs.
- Your product has shared credits, concurrent requests, or agents that can consume an organization's balance without a predictable final cost.
- You sell AI workloads to enterprises and need budgets, spend caps, usage attribution, or visibility across teams and models.

## Related talks

- [Mastering AI Pricing](https://aietalks.com/talks/mastering-ai-pricing) (Mayank Pant, Stripe, 24:19)
- [The Price of Intelligence - AI Agent Pricing in 2025](https://aietalks.com/talks/the-price-of-intelligence-ai-agent-pricing-in-2025) (Chz, Orb, 20:38)
- [Revenue Engineering: How to Price (and Reprice) Your AI Product](https://aietalks.com/talks/revenue-engineering-how-to-price-and-reprice-your-ai-product) (Kshitij Grover, Orb, 15:39)
- [When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS](https://aietalks.com/talks/when-ai-agents-pay-and-sellers-monetize-building-x402-apps-on-aws) (Anil Nadiminti, AWS, 20:41)
- [Why Off-the-Shelf AI Doesn't Understand Money](https://aietalks.com/talks/why-off-the-shelf-ai-doesnt-understand-money) (Udi Menkes, Intuit, 19:50)
