FinanceBench

11.4% cheaper on 10-K filings, all in.

FinanceBench is the public 10-K question set from Patronus AI: dense filings where any page can hold the answer. Needlepath cut the bill 11.4% with our fee included, delivered 64.7% fewer input tokens, kept 81% of full-context accuracy, and sent the whole filing whenever it was not sure. 70 questions, the same answer model and the same scorer for every arm.

64.7%
fewer input tokens across the whole corpus
81%
of full-context accuracy
11.4%
lower bill, our fee included
The Field

Superior to full context on cost and tokens, at four fifths of its accuracy.

Accuracy 81% of full context Fewer tokens 64.7% fewer Less model spend 63.2% less Cost advantage 11.4% lower Fewer tokens per correct answer 56.4% fewer
Needlepath Full context Hosted compression service

Needlepath answered 47 of 70 questions correctly with 64.7% fewer input tokens and 11.4% off the bill. The complete filing answered 58 of 70 at every token. A hosted compression service kept more of the accuracy, 54 of 70, and reads and charges for every token: all in, it cost 19.2% more than sending the whole filing.

Accuracy: share of the 70 questions answered correctly (67.14% Needlepath, 82.86% full context, 77.14% hosted compression). Tokens delivered to the answer model across the run, whole-filing calls included (1.48M, 4.19M, 1.99M). Answer-model spend before any fee ($0.197, $0.536, $0.261). Bill with fees ($0.475, $0.536, $0.639; the compression service's fee is $0.378 at $0.10 per 1M processed tokens, rate checked 2026-08-14; full context pays no fee). Tokens per correct answer (31.5K, 72.2K, 36.8K). Against full context the compression service removed 52.5% of input tokens, spent 51.3% less on the answer model, and used 49.0% fewer tokens per correct answer, at 19.2% more all in. Needlepath (r4) and full context measured August 2026, the same model at temperature 0, the same scorer; the hosted compression service in its own run on the same items. A local compression model finished 4 of 70 filings inside its two-minute limit and is not drawn. Harness: context-selection-bench.

The Tokens

91.4% fewer tokens on the calls it cut.

Needlepath cut 91.4% of input tokens on 47 of 70 calls. Across the whole corpus it delivered 1.48M tokens instead of 4.19M in whole filings: 64.7% lower. Original pages, nothing rewritten.

The Fail-Safe

Nothing gets missed.

When it is not sure, Needlepath sends the whole filing. On this run that was 23 of 70 calls, and on each of them the model saw exactly what it would have seen without Needlepath.

The Bill

11.4% lower at the baseline.

Working with Needlepath saved 11.4% on this run, our fee included: $0.475 against $0.536. With your agent keeping context history locally, those savings may compound with each call.

Needlepath fee at the live rate card ($0.09 per 1M metered input tokens) on the calls it reduced. Full context pays no fee.

Fit

Which shape is your traffic?

Dense filings trade a share of accuracy for a much smaller bill; you set the weights. Long records with a short answer inside give you both at once: on 2,600 RULER questions Needlepath answered +12.9 pp better and cut the bill 44.31%, all in. The RULER results show it in full. A workload review tells you which shape your traffic is.

Start

Measure it on your own corpus.

Keep your model. Change the endpoint, run your own traffic, and compare the answers and the bill. Pay as you go includes 111M tokens.

Provenance

Benchmark. FinanceBench by Patronus AI: Islam et al., FinanceBench: A New Benchmark for Financial Question Answering (2023), github.com/patronus-ai/financebench, CC BY-NC 4.0. We used the publicly available test set for evaluation only, with Patronus AI's written permission, and no Needlepath component was trained on it.

Run. The 70 questions whose filings fit under the service input ceiling, August 2026, on the deployed point r4. Tokens on the answer model's own count. Harness, item list and scoring rules: context-selection-bench.