Stop Ordering AI Takeout: A Cookbook for Winning When You Build In House

Jan Siml10:45 · Jun 2025 · 345 views
Thumbnail for Stop Ordering AI Takeout: A Cookbook for Winning When You Build In House Watch on YouTube
TL;DR
  1. 1

    A small in-house team used two developers and roughly 10 sprint weeks to build a system that produced several million dollars of ARR.

  2. 2

    The team focused on one sales-alert workflow, measured the path to revenue, and pushed daily digests instead of waiting for users to open a chat interface.

  3. 3

    Better data, user feedback, integrations, and workflow design produced more value than switching between larger and smaller models.

Summary

Jan Siml argues that companies should build AI in-house when they own the data, understand the workflow, and can work closely with users. His team started with a narrow sales-alert use case and went deep instead of building a broad agent system or a large evaluation suite. Two developers spent roughly 10 sprint weeks on the system, which led to several million dollars of ARR, although Siml cannot share the exact figure. The team linked each AI action to a revenue outcome, sent users proactive daily digests, and used a chat interface only for unexpected tasks. Siml says saved time has value only when the system guides people toward higher-value work. He also argues that better data and more workflow triggers beat expensive model upgrades. The talk is a practical case for small, focused systems whose success is measured by business outcomes and shaped by user feedback.

Key ideas
00:52

A small in-house build produced several million dollars of ARR

In Q1 2024, Siml's team faced a build-or-buy decision. They chose to build with two developers and roughly 10 sprint weeks of effort. The resulting system produced several million dollars of ARR and earned a group-level award. Siml cannot share the exact revenue figure because of legal limits, but says it was large enough to end the question of why the company should build. He contrasts this with a recipe built around giant evaluations, multi-agent systems, and expensive models, which he says would have delayed the launch and cost too much for the team's internal needs.

01:53

Build when the workflow, data, and users are already yours

Siml frames generic software as a hotel buffet and an internal system as a home kitchen. Buying makes sense when a company needs vendor integrations or best practices from other industries. Building has an advantage when the company already owns the data, knows the exact steps colleagues take to close deals, and can involve those colleagues in checking outputs. The team also sat close to its users, so changes could ship quickly and the interface could use language familiar to them. Since the infrastructure was mostly already paid for, the marginal cost was very low. His rule is to explore the unknown by buying, then build in-house once the workflow belongs to the company.

03:27

One painful job with a clear value event is a better starting point

The first lesson is to go deep on one job to be done. Siml says an in-house team can focus on a painful task without chasing a large total addressable market. The task should have a clear value event, meaning the dollar-based outcome that justifies the work. His team began with a simple sales-alert use case and expanded by asking what the alerts were for and what else users needed to do. Those answers came from conversations with users. Staying focused also kept the system simple and let the team avoid adding agentic behavior before it was needed.

04:25

Revenue funnels matter more than offline evaluation scores

Siml says offline evaluations do not sign contracts, and board members ask whether revenue moved rather than asking for an F1 score or NDCG. He still considers evaluations useful, but compares them to smoke alarms. The team instrumented the full path from the AI task to the dollar outcome, so it could identify the money associated with an action. Once the system was linked to revenue, prioritization became easier because ideas could be discussed in terms of expected sales. Siml also recommends automating team performance reports and preparing leaderboards, which can create competition, involve leadership, and reveal people struggling with new workflows.

06:19

Proactive daily digests beat waiting for users to open the interface

After measuring revenue, the team stopped waiting for users to ask for help. Siml says the best interface is one users do not need to open because the system already knows the next business step. His team built a motion that sent daily digests with what users needed to know that day. The chat interface remained available as a fallback for unexpected and unplanned tasks. This approach moved the product from a place users had to visit toward a system that inserted useful information into their existing work.

06:48

Time saved has value only when the system guides the next action

Siml says an AI system should guide action rather than only provide information. Saving 30 minutes has little value if a user fills that time with low-value email work. The system needs to turn freed-up minutes into higher-value activity, based on what the business knows matters most. As the team built more revenue funnels, it learned where to direct users' spare time and attention. The proactive system surfaced actions users might not have thought to take. Siml says it achieved 20 points higher NPS and an order-of-magnitude higher engagement than the chat application.

08:08

Better data and workflow triggers beat expensive model upgrades

When deciding where to spend limited development time, Siml recommends investing in data and workflow depth before larger models. He says o3 was 60 times more expensive and an order of magnitude slower than 4.1 mini, with the main production effect showing up in cost. The team's strongest results came from adding more triggers to alert users and going deeper into what they needed. Switching between the normal and mini model series changed costs and evaluations, but did not change the user-facing results. Siml's advice is to build for user needs rather than for experiments the team wants to try.

09:07

User feedback creates a revenue flywheel

When the team focused on what users valued instead of model benchmarks, users began providing ideas for improvements because the feedback loop felt close to their work. The team could run experiments based on those suggestions, which increased adoption and generated more data. That data improved prioritization and produced further ideas. Siml describes this as a revenue flywheel. His final recipe is to focus on one painful job with clear dollar value, track the outcome to revenue, push insights proactively, direct saved time toward valuable work, and invest in basic data and workflow improvements.

"Saving 30 minutes is worthless if users just fill it with an email sludge."07:12
Who should watch
  • You are deciding whether to buy a general AI product or build around a workflow your company already understands.
  • Your team has an internal assistant or chat app, but you cannot connect its usage to revenue or other business outcomes.
  • You are spending time on larger models and evaluation suites while users still need better data, triggers, and guidance for their next action.