We make your AI governance blazing fast.
Runtime context governance that keeps up with your business.
Apply hundreds of policies on large contexts
Maximize capacityGovern agent context in milliseconds
Increase speedLower your governance costs
Lower costsOne call in front of every governance check or decision.
QuestionWhat is Adobe's year-over-year change in operating income from FY2015 to FY2016?
- Reads 474 tokens, not 46,752
- Checks in 0.23 s instead of 10.5 s
Measured across the guardrail run, 960 calls.
Kept the one page that answers it, Adobe's 2016 income statement. The 125-rule policy check then ran in 0.23 s instead of 10.5 s.
Recorded call in front of a commercial AI guardrail with a 125-rule policy. Question from FinanceBench (Patronus AI); pages from public 10-K filings. Times are the median of three runs. See real recorded calls
Up and running in minutes. Under ten lines of code.
SaaS
Call our API from your stack in under ten lines of code.
from langchain.agents import create_agent
from needlepath_langchain import NeedlepathMiddleware
# NEEDLEPATH_API_KEY is read from the environment. Nothing else in the agent changes.
agent = create_agent(model, tools, middleware=[NeedlepathMiddleware(operating_point="np-2026-08-r4")])
answer = agent.invoke({"messages": [("user", "Which Corvane Freight purchase orders are still open past their delivery date?")]})import os
from needlepath import NeedlepathClient, ContextRecord, TaskSpec
from openai import OpenAI
client = NeedlepathClient(api_key=os.environ["NEEDLEPATH_API_KEY"], operating_point="np-2026-08-r4")
# One record per item the agent holds: purchase orders, the supplier contract, the conversation so far.
records = [ContextRecord(text=po.text, kind="external_data", id=po.id, title=po.title) for po in purchase_orders]
prompt = "Which Corvane Freight purchase orders are still open past their delivery date?"
result = client.select(records=records, task=TaskSpec(prompt=prompt), max_context_tokens=4000)
context = result.rendered_context if result.applied else "\n\n".join(r.text for r in records)
# Your model call, unchanged, with the context ahead of the prompt.
answer = OpenAI().chat.completions.create(
model="your-model", messages=[{"role": "user", "content": f"{context}\n\n{prompt}"}]
)import { NeedlepathClient } from "@nextmoca/needlepath-sdk";
import OpenAI from "openai";
const client = new NeedlepathClient({ apiKey: process.env.NEEDLEPATH_API_KEY!, operatingPoint: "np-2026-08-r4" });
// One record per item the agent holds: purchase orders, the supplier contract, the conversation so far.
const records = purchaseOrders.map((po) => ({ text: po.text, kind: "external_data", id: po.id, title: po.title }));
const prompt = "Which Corvane Freight purchase orders are still open past their delivery date?";
const result = await client.select({ records, task: { prompt }, maxContextTokens: 4000 });
const context = result.applied ? result.response!.renderedContext : records.map((r) => r.text).join("\n\n");
// Your model call, unchanged, with the context ahead of the prompt.
const answer = await new OpenAI().chat.completions.create({
model: "your-model", messages: [{ role: "user", content: `${context}\n\n${prompt}` }],
});# config.yaml: every client behind the proxy gets selection, with no client-side change.
guardrails:
- guardrail_name: "needlepath"
litellm_params:
guardrail: needlepath_litellm.NeedlepathGuardrail
mode: "pre_call"
default_on: true
operating_point: "np-2026-08-r4"
history_max_tokens: 8000
preserve_recent: 2
# export NEEDLEPATH_API_KEY="np_live_..." in the proxy's environment, then: litellm --config config.yaml# One record per item the agent holds. Use rendered_context from the response ahead of the prompt.
curl https://api.nextmoca.com/v1/context/select \
-H "Authorization: Bearer $NEEDLEPATH_API_KEY" -H "Content-Type: application/json" \
-d '{
"records": [
{ "id": "po-2026-0412", "kind": "external_data", "title": "PO 2026-0412, Corvane Freight", "text": "..." },
{ "id": "contract-corvane", "kind": "external_data", "title": "Corvane Freight delivery terms", "text": "..." }
],
"task": { "prompt": "Which Corvane Freight purchase orders are still open past their delivery date?" },
"budget": { "max_context_tokens": 4000, "operating_point": "np-2026-08-r4" },
"render": true
}'Four answers before you decide.
What does Needlepath do?
Needlepath decides what enters each model call. Send the records your agent holds with the question the model is about to answer, and it returns the ones the step needs, unrewritten, in milliseconds. When the whole set is what the question needs, the whole set goes through unchanged.
How much less context does a call carry?
On the NVIDIA RULER benchmark, 2,600 questions, 52.9% fewer input tokens were sent, counting the 24.0% of calls that received the complete document. Answers were +12.9 pp better than sending everything.
Does it rewrite or summarize my documents?
No. Needlepath sends your own records, unrewritten, and only the ones that answer the question. It does not summarize, compress or rerank the content of a record.
What happens when the whole file is what the question needs?
It hands over the whole file. When a smaller context cannot be shown to be safe, the complete document goes through unchanged, so what you risk on that call is paying full price.
Still deciding? Talk to the founders, or start with an API key.