Measured

Measured on public benchmarks, with the receipts.

Scored on the NVIDIA RULER benchmark, 2,600 questions, one run, official scorer, against sending everything.

  1. 1Answers were +12.9 pp better than sending the full document.
  2. 2Answering was 504 ms faster at p95 with Needlepath in the stack, end to end. The decision itself took 25.0 ms on average.
  3. 352.9% fewer input tokens were sent, counting the 24.0% of calls that received the complete document.
Other preparation methods, measured on the same items, took 1.13 to 32.4 s per item in selection latency alone, some needing a GPU.

Every step reads everything, whether or not it should.

Each tool call, document, memory read and workflow step adds to the context. The next check reads all of it: more tokens, more time, and more chances to act on the wrong detail. The usual fixes each cost something too.

SEND EVERYTHING

Attention gets diluted

Every extra page costs input tokens and competes for the model's attention, whether or not it helps this step.

REWRITE IT SHORTER

Details go missing

A summary fits nicely, and can leave out the exact name, amount or order detail the next step needs, without saying so.

BUILD RETRIEVAL

You own a system

Chunking, tuning and evals that drift every time the corpus changes, on someone's on-call rotation.

Needlepath sends your own records, unrewritten, and only the ones the step needs. When the whole file is what the moment needs, it hands over the whole file, with no second model in the loop, and a receipt listing what was sent and what was left out.

The same public exam, task by task.

Each measured against sending everything.

WINFinding one buried fact+17 points at 8K, +21 at 16KStatistically significant at both sizes.
WINMany needles, one answer+16 points at 8K, +20 at 16KA consistent gain.
WINCounting and extraction+15 points at 8K, +23 at 16KTask-level gains at both context lengths.
WINHard multi-key lookups+13 points at 8K, +12 at 16KStatistically significant.
EVENLong-document QApooled -2.0 points, interval includes zeroStatistically even with sending everything. Needlepath keeps more, or steps aside, where coverage matters.

Run the same measurement on your own workload.

What it means for my business

Every figure on this page is from one run of Needlepath at operating point np-2026-08-r4 on RULER, 2,600 questions, official scorer. Full results.