Small Models, Big Results: Training a Finance Agent for Under $500

Charles Dickens, Snorkel AI17:16 · Oct 2026 · 4,637 views
Thumbnail for Small Models, Big Results: Training a Finance Agent for Under $500 Watch on YouTube
TL;DR
  1. 1

    A specialized Qwen3 4B model trained with reinforcement learning reached almost 60% accuracy on FinQA, compared with 51% for a 235B model.

  2. 2

    The FinQA dataset uses SEC 10-K tables, expert-designed question types, programmatic checks, independent agent reviews, and manual financial review.

  3. 3

    Simple training data and a binary pass/fail reward worked better than multi-table curricula and detailed intermediate-step rewards because tool-use discipline was the main limitation.

Summary

Charles Dickens describes a collaboration between Snorkel AI and UC Berkeley's Sky Computing Lab to train a finance agent with a Qwen3 4B model. The team built FinQA from SEC 10-K filings, converting roughly 6,900 tables into verified question-answer examples. Models were tested in an environment where an agent queried tables through tools and received a binary correctness reward. After reinforcement learning with rLLM, the 4B model reached almost 60% on a held-out benchmark of 290 questions, beating a 235B model at 51%. The model also improved on harder questions involving two to five tables, without losing general tool-calling ability. Dickens argues that the main problem was disciplined tool use, including avoiding invented schemas, oversized queries, and repeated failed strategies. Simpler training data and rewards produced the strongest results. He closes with an evaluation framework based on environment complexity, autonomy horizon, and output complexity.

Key ideas
00:36

Specialization can matter more than parameter scale for enterprise agents

Dickens argues that enterprise workloads often value reliability and specialization more than raw parameter count. Financial agents must work with complex schemas, legacy APIs, technical debt, and long workflows without compounding errors. They also need to be auditable. In this setting, a model trained on data that matches the target domain and tasks can reach parity with a frontier model at a fraction of the cost. He compares the choice to calling a tax specialist who knows the forms, tools, and rules instead of a general polymath.

04:56

FinQA turns SEC filings into verified tool-use tasks

The team built FinQA from annual 10-K reports required by the SEC. They pulled filings from the Edgar system and used Qwen3 30B to produce about 6,900 SQL tables. Each table generated a question-answer pair, with a question taxonomy developed with financial experts and metadata such as table names, columns, and lineage. The resulting tasks require planning, tool calls, and reasoning, while each answer has one verifiable final result. The split contains about 4,000 training examples, 500 validation examples, and 290 held-out benchmark examples, with no company overlap between splits.

06:03

Three verification layers reduce errors in generated financial data

FinQA uses programmatic consistency checks to compare generated content against metadata, including lineage, table names, and columns. Independent agents then review the examples automatically. Financial experts perform the final manual review to check that the questions and answers are realistic. This process addresses the risk that generated tables or questions contain invented schemas. Dickens presents the resulting dataset as expert-validated data for both training and evaluation.

06:48

Frontier models failed through poor tool discipline

The team found recurring failures even in frontier models. Models hallucinated tables and column names, sometimes because those structures resembled patterns from their pretraining data. They also flooded their own context with badly planned queries, such as using SELECT STAR, which could return too much information. A further problem was poor recovery: the 235B model sometimes repeated the same failed strategy instead of using error messages to change its approach. Dickens says these failures also appeared in insurance underwriting.

08:16

RL in rLLM trained the 4B model for less than $500

The team used the open-source rLLM framework, which works with agent frameworks through a decorator pattern and supports multiple reinforcement-learning algorithms and training backends. The agent ran in a ReAct loop and queried the generated tables. A language-model judge, GPT-5 nano, assigned a binary correctness reward using reference-based evaluation. The setup used Qwen3 4B, GRPO, about 1,000 concurrent environments, and eight H100 GPUs. Dickens reports roughly $420 in compute costs and $40 for the judge API.

10:20

The trained 4B model beat the 235B model on held-out finance questions

On 290 expert-curated held-out examples, the trained 4B model reached almost 60% pass rate, while the 235B model reached 51%. Dickens says the base model's accuracy more than doubled after training. The comparison tested models on financial questions that required interacting with specialized tables and tools. The result supports his claim that a smaller model can outperform a much larger general model when the task and environment are narrowly defined.

10:55

Tool-use skills transferred to multi-table questions and general tool calling

The team tested harder FinQA reasoning examples involving two to five tables, sequential decisions, and planning. Training only on simpler single-table examples still improved performance on the multi-table variant. The team also evaluated the model on BFCL, a benchmark for general tool calling. Overall accuracy improved slightly, with minor gains on multi-turn and memory tasks. Dickens says the finance specialization did not erode the model's broader tool-use competence after RL fine-tuning.

12:35

Simple data and a binary reward beat more elaborate training designs

Ablation studies compared training on FinQA, training on single- and multi-table data, and a curriculum that moved from single-table to multi-table examples. Training only on the simpler set produced the most lift. The team also tested a detailed rubric that rewarded intermediate steps, table access, and query completeness. That approach underperformed a single binary reward. Dickens concludes that the bottleneck was tool-use reliability rather than reasoning depth. Once the model learned the basic tool-use behavior, it could compose those skills on harder tasks.

14:09

Agent evaluation needs realistic environments and longer decision horizons

Dickens proposes evaluating agents along three axes. Environment complexity measures how realistic, dynamic, and difficult the agent's working environment is. Autonomy horizon measures whether the evaluation covers the length of work people expect agents to perform and whether the agent makes safe decisions over that period. Output complexity measures the range of outputs required in day-to-day work. He says current evaluations still focus heavily on text artifacts, leaving room for more rigorous tests of agents operating in realistic settings.

"For enterprise AI workloads, reliability and specialization can actually outweigh raw parameter scale."00:36
Who should watch
  • You are choosing between a large general model and a smaller model for a narrow enterprise workflow, and you want evidence about the tradeoff.
  • You are building an agent that must query structured data and recover from tool errors rather than only produce text.
  • You need an evaluation plan for agents that covers realistic environments, long autonomous runs, and outputs beyond chat responses.