Half the tokens. Better answers.
One public exam, five ways to prepare the packet, the same questions and the same scorer for each. 13 RULER tasks, 100 questions each, at 8K and 16K context, 2,600 questions in all.
One method beat the baseline.
Needlepath answered 55.6% of the questions correctly against 42.7% for the complete file. The gap is +12.9 pp, with a 95% confidence interval of +11.5 to +14.3 and p = 2.1e-24.
Accuracy change against sending the complete file, in percentage points. All arms: the same 2,600 questions, the same model at temperature 0, the same scorer. Needlepath (r4) and full context measured September 2026; the other arms in their own runs on the same items. Reproduce it with the published harness at context-selection-bench.
When a smaller context cannot be shown to be safe, the whole file goes through.
On 24.0% of these calls, 625 of 2,600 rows, Needlepath sent the complete file rather than a smaller context it could not stand behind. Those calls are counted inside the 52.9%: 14.45M input tokens delivered across the whole run against 30.64M sending the complete file every time. On those calls the model sees exactly what it would have seen without Needlepath, and you pay what you would have paid anyway. The saving is net of them.
Cheaper per correct answer, after our fee.
A correct answer cost $0.0266 through Needlepath against $0.0621 sending the complete file, which is 2.3x as many correct answers per dollar. That figure is all in: it already includes what Needlepath charges. Before our fee, the answer model alone cost $0.0200 against $0.0552, which is 2.8x: that figure excludes what Needlepath charges, so it is not the one to plan against. Our fee takes back part of the saving, and the 2.3x is what is left for you. Across the same 2,600 answers the whole bill was $38.44 against $69.02, 44.31% lower.
Needlepath fee at the live rate card ($0.09 per 1M metered input tokens); requests that return every record unchanged are not charged.
The decision costs milliseconds, not seconds.
Selection took 25.0 ms on average per call, measured in process across this run: 16.0 ms at the median, 71.8 ms at the 95th percentile, 262 ms at the slowest call in the run. No model runs in the selection path, so there is no second deployment to keep warm. This is the selection decision only, not the time your whole request takes.
Where selection has nothing to remove.
RULER is a long haystack with a short answer inside it, which is the shape where selection pays. On short prompts, on context you have already trimmed, on work where the output dominates the bill, and on files where nearly every line can be the line that matters, there is little to take out and the complete file is the right packet. Dense filings are measured too: the FinanceBench results, two thirds fewer tokens on 10-K filings.
Point your calls at the endpoint.
Keep your model. Change the endpoint, run your own traffic, and compare the answers. Pay as you go includes 111M tokens and signup takes no card.