Measured on public benchmarks, with the receipts.
Scored on the NVIDIA RULER benchmark, 2,600 questions, one run, official scorer, against sending everything.
- 1Answers were +12.9 pp better than sending the full document.
- 2Answering was 504 ms faster at p95 with Needlepath in the stack, end to end. The decision itself took 25.0 ms on average.
- 352.9% fewer input tokens were sent, counting the 24.0% of calls that received the complete document.
Every step reads everything, whether or not it should.
Each tool call, document, memory read and workflow step adds to the context. The next check reads all of it: more tokens, more time, and more chances to act on the wrong detail. The usual fixes each cost something too.
Attention gets diluted
Every extra page costs input tokens and competes for the model's attention, whether or not it helps this step.
Details go missing
A summary fits nicely, and can leave out the exact name, amount or order detail the next step needs, without saying so.
You own a system
Chunking, tuning and evals that drift every time the corpus changes, on someone's on-call rotation.
Needlepath sends your own records, unrewritten, and only the ones the step needs. When the whole file is what the moment needs, it hands over the whole file, with no second model in the loop, and a receipt listing what was sent and what was left out.
The same public exam, task by task.
Each measured against sending everything.