Half the tokens. Better answers.
One public exam, five ways to prepare the packet, the same questions and the same scorer for each. 13 RULER tasks, 100 questions each, at 8K and 16K context, 2,600 questions in all.
One method beat the baseline.
Needlepath answered 55.6% of the questions correctly against 42.7% for the complete file. The gap is +12.9 pp, with a 95% confidence interval of +11.5 to +14.3 and p = 2.1e-24.
Accuracy change against sending the complete file, in percentage points. All arms: the same 2,600 questions, the same model at temperature 0, the same scorer. Needlepath (r4) and full context measured September 2026; the other arms in their own runs on the same items. Reproduce it with the published harness at context-selection-bench.
When it cannot prove it, it steps aside.
On 24.0% of these calls, 625 of 2,600 rows, a smaller context could not be shown to be safe, so Needlepath sent the complete file. Those calls are inside the token count, not excluded from it: 14.45M input tokens delivered across the whole run against 30.64M sending the complete file every time, which is the 52.9% above. What you risk is paying full price, never a wrong answer.
Cheaper per correct answer, after our fee.
A correct answer cost $0.0271 through Needlepath against $0.0621 sending the complete file, which is 2.3x as many correct answers per dollar. That figure is all in: it already includes what Needlepath charges. Before our fee, the answer model alone cost $0.0200 against $0.0552, which is 2.8x: that figure excludes what Needlepath charges, so it is not the one to plan against. Our fee takes back part of the saving, and the 2.3x is what is left for you. Across the same 2,600 answers the whole bill was $39.15 against $69.02, 43% lower.
The decision costs milliseconds, not seconds.
Selection took 25.0 ms on average per call, measured in process across this run: 16.0 ms at the median, 71.8 ms at the 95th percentile, 262 ms at the slowest call in the run. No model runs in the selection path, so there is no second deployment to keep warm. This is the selection decision only, not the time your whole request takes.
Where selection has nothing to remove.
RULER is a long haystack with a short answer inside it, which is the shape where selection pays. On short prompts, on context you have already trimmed, on work where the output dominates the bill, and on files where nearly every line can be the line that matters, there is little to take out and the complete file is the right packet.
Point your calls at the endpoint.
Keep your model. Change the endpoint, run your own traffic, and compare the answers. Pay as you go includes 111M tokens and signup takes no card.