Stop paying for stale chat history and making your AI responses worse.
AI agents pile up state: documents, tool results, approvals, history. Sending it all makes answers worse and bills bigger, and rewriting it with another AI is slow and loses details. Needlepath selects what matters before every model call... in milliseconds, with results proven on public, industry-standard tests.
Every model call gets a briefing. The question is who prepares it.
Before every call, something decides what the model reads. There are four ways to prepare that packet:
Send the complete file
The answer is in there, but buried, and you pay for every page.
Rewrite it shorter
A summary fits nicely, but exact names, numbers, and order details can vanish in the rewrite.
Heavyweight selection
Picking pages beats rewriting them, but this needs its own expensive machine and takes seconds per question.
A focused packet of originals
Original records, untouched, selected in milliseconds. And when the whole file is what the moment needs, it hands over the whole file.
Better answers than sending everything.
One matched public exam (RULER), five ways to prepare the packet, two context lengths. Higher is better.
Method: same questions, same model family, official scoring paths, paired statistics. Each method ran at its own recommended settings. Per-item results and reproducibility artifacts are published for inspection.
Suites: RULER · BFCL · SQuAD v2. Scoring: the benchmarks' official code, no house grading. Statistics: McNemar + bootstrap CIs. Artifacts: sha256 manifests.
The same public exam, task by task.
Per-task results from the flagship public run. Every number carries its own statistical label, and we never blend them into a composite score.
Finding one buried fact
+12 points at 8K, +15 at 16K
Statistically significant at both sizes.
Many needles, one answer
+13 points at both sizes
A consistent gain.
Counting and extraction
+7 to +18 points across sizes
Task-level gains at both context lengths.
Hard multi-key lookups
+12 points at 8K
Statistically significant.
Long-document QA
Statistically even with sending everything, within a few points in both directions.
Conservative by design
Keeps more or steps aside where order, forms, or answerability need coverage.
Four contests. The same public exams. Official scorers.
All four ran on public, industry-standard tests, scored by the official scoring code.
When the decisive fact is buried
The critical detail sits deep inside a 100-message thread. Sending everything put the answer in the room, surrounded by noise that cost money and competed for the model's attention. The rewrite lost exact wording and scored below sending everything.
When the answer cannot wait
A live conversation, where every pause is visible. The heavyweight systems took one to over thirty seconds to decide what to keep.
When every field must remain exact
Orders, forms, controlled documents. A summary can be almost right and operationally wrong: a dropped decimal breaks the order.
When the complete file is the right packet
Some work needs the entire history. Needlepath is the only contestant with a tested gate for exactly this: it detects those workloads and hands over exactly what you would have sent anyway, at no extra cost.
Where Needlepath acts, and where it stands aside.
Three disclosed cases, all in the public results. Honesty is the feature no compressor ships.
Clean single-document inputs: nothing to select, so Needlepath does nothing, by design.
Distractors inside a single record: Needlepath selects among records, it does not edit within them.
Agent loops where the whole history is load-bearing: the gate detects this per workload and stands aside, handing over exactly what you would have sent anyway, at no extra cost.
Stood aside on 100% of steps; the model received the full history unchanged.
Zero changed answers across configurations.
Milliseconds inside the loop.
The context decision, measured per item on the public benchmark, excluding model time.
A selection layer, not a compressor.
Parse the objective
Evaluate the immediate request: task intent, tool intent, output type, entities.
Calculate relevance
Fast filters identify the records most relevant to the current task, in 8.5 to 14.5 ms per decision, with no model call.
Route the state
Ship the relevant records intact, preserve required evidence exactly, and pass the full context through when the workload calls for it.
POST /v1/context/select { "task": { "query": "resolve the customer's refund request", "intent": "tool_call" }, "records": [ ...documents, tool results, history... ], "constraints": { "max_selected_tokens": 8000 }, "operating_point": "np-2026-07-r1", // immutable, versioned label "mode": "shadow" // evaluate without changing what you send }
In our cloud, or inside yours.
Start hosted in an afternoon, or keep everything inside your network. Same engine, same results, your call.
Drop in. One API key.
Install the SDK, point it at your workload, and Needlepath prepares the context before every model call. Start in shadow mode: watch the savings before changing anything you send.
- TypeScript and Python clients
- Fast and intelligent context preparation
- Shadow mode first: zero-risk evaluation
- Transparent pricing
Runs where your data lives.
Deployed into your VPC or data center, tuned for your workloads. Context never leaves your network, and the selection path sits next to your models.
- Isolated deployment in your cloud or data center
- Higher throughput with lower latency
- Dedicated support
Hyperscale operators and engineers who built and shipped agentic infrastructure at billion-dollar revenue scale. We are building Next Moca because we know what the next ten years of enterprise AI actually need.

Repeat builder. Co-founded and served as CTO of Green Piñata Toys (acquired). Director of Engineering at Oracle (OCI) and Brightcove, with prior leadership roles at Juniper, Cisco, and others. Built and scaled world-scale services for customers like ByteDance and HBO. Architect behind Next Moca's agentic orchestration and infrastructure layers. MBA (Babson), M.S. Computer Science (Penn State), B.E. Computer Engineering (University of Mumbai).

Product-driven inventor with a track record of bootstrapping new products, blending deep engineering with go-to-market focus. Director of Product at Adobe, with prior leadership roles at Microsoft, Oracle, and others. Holds multiple patents in distributed systems. Architect behind Next Moca's agentic foundations and core AI services. Executive MBA (Michigan Ross), M.S. Computer Science (USC), B.E. Computer Engineering (University of Mumbai).

Engineering and UX leader with 25+ years at Sun Microsystems and Oracle. Founding Member of Technical Staff at Next Moca, leading Agent Experiences. Recognized expertise in UI/UX architecture, data security, and full-stack engineering. Building the intuitive, secure interfaces that bridge human creativity with intelligent automation.
Operators who built and shipped enterprise platforms at billion-dollar scale. They're betting on Next Moca because they've lived the problem we're solving.

Ran product and strategy at Brightcove (CPO) and Ellucian (CPO) over two decades of enterprise SaaS leadership. Now Managing Partner at ND Labs and lead investor in Next Moca.

Head of ML Platform at Capital One, operating production AI infrastructure for one of the largest regulated U.S. banks. Ph.D. in EECS from UC Berkeley, M.S. from UIUC, and B.Tech. from IIT Bombay.
Make every token earn its place.
Pick one high-volume workflow. Replay your own traces, with and without Needlepath, same model. Read the report: answer quality, tokens removed, cost saved, and where it chose to stand aside.
Thanks. Expect us to get in touch soon.
In the meantime, you can also reach us directly at kiran@nextmoca.com or swanand@nextmoca.com.
kiran@nextmoca.com · swanand@nextmoca.com · Boston, MA · Palo Alto, CA