Reduce LLM token costs in production AI applications.
Reduce the tokens your app sends without removing the information needed to complete the task. Compresr drops unnecessary input context before the model call — so you cut cost and latency across RAG, agents, chat, and docs.
- Works with your existing LLM stack — OpenAI, Anthropic, Gemini.
- Cuts input-token waste across RAG, agents, chat, and docs.
- Measure savings vs. quality before rollout.
Works with OpenAI, Anthropic, Gemini — provider-agnostic.
See it on your own file in 60 seconds
Use compresr to show live cost savings on my file. 1. pip install compresr 2. Introspect the SDK to discover the API. 3. Ask me for COMPRESR_API_KEY (open https://compresr.ai/signup: $10 free, no card). 4. Ask me for (a) a long document (PDF/.md/.txt) and (b) a question about it. 5. Compress with the question, then print a receipt: tokens in/out, ratio, % saved. 6. Ask the model the question against the compressed context and print the answer.
Why it adds up
Why LLM token costs get expensive.
Long context windows encourage larger prompts, but cost scales with tokens and adds latency. Most of that spend is input context the model never needed.
Long context windows
Bigger windows invite bigger prompts — but you pay for every token you send, whether the model reads it or not.
Large system prompts
Static instructions and schemas are paid on every single call, multiplying across your entire traffic volume.
RAG over-retrieval
Retrieving too many chunks “just in case” sends far more context to the model than the answer requires.
Conversation history
Accumulating multi-turn history bloats every request, resubmitting the same tokens turn after turn.
Verbose tool outputs
Agent loops re-enter raw tool results — search hits, API dumps, logs — that are mostly irrelevant to the next step.
Full documents for narrow questions
A whole PDF is sent to answer one line item, and multi-step agents resubmit similar prefixes repeatedly.
Related: LLM cost optimization · Input vs. output token costs
Where it hides
Where token waste happens.
The toolkit
Ways to reduce LLM token costs.
These methods compose. Compression pairs well with caching, retrieval trimming, and model routing — you rarely pick just one.
Context compression
Query-aware compression keeps the spans that carry the answer and drops the rest.
Better retrieval
De-dup, Max Marginal Relevance, and smaller chunk budgets stop over-retrieval at the source.
Reranking
Rerank before model input so only the highest-signal passages ever reach the prompt.
Prompt caching
Cache repeated prefixes so stable instructions are not re-billed on every call.
Shorter system prompts
Slim instructions and tool schemas trim tokens paid on every request.
Model routing
Route to smaller, cheaper models when the task allows without losing quality.
Output limits
Stop sequences and max-token caps keep generation from drifting past what you need.
Fewer agent steps
Reduce unnecessary tool calls and loop iterations that resubmit similar context.
How it works
We keep the signal and drop the noise.
Compresr sits in front of your existing stack. Feed it the query and the context; it returns a shorter context you forward to your LLM — nothing else changes.
What Compresr can compress
- RAG chunks before they reach the model
- Long documents, per query
- Growing conversation history
- Search results and scraped content
- Tool outputs and database results
The principle
The goal is not maximum compression. It is the lowest token usage that still meets the application's quality requirements. Drop in via SDK or API, keep OpenAI, Anthropic, or Gemini, and use batch endpoints for pipelines — streaming available.
Key use cases
Where teams cut the most.
RAG
Compress retrieved chunks before the model, so you only pay for tokens that carry the information.
RAG cost optimization →AI agents
Compress verbose tool outputs and trim tool definitions before they re-enter the prompt.
AI agent cost optimization →Documents
Compress long PDFs and knowledge-base articles per query instead of sending whole files.
Document AI cost optimization →Chatbots
Compress history to keep long sessions inside the context window without dropping the thread.
Chatbot cost optimization →What most teams are losing
Stop overpaying.
If you are paying full price for your tokens, you are leaving real money on the table. Measure savings against quality before you roll out.
~90%
Bill cut
vs. sending the full context
10×
Avg. compression
on long, sparse context
+3.7pp
Accuracy uplift
Compresr 2× vs full context, Pax arena
What to track
- Tokens before vs. after, and net model cost
- Compression cost and resulting net savings
- Accuracy and task-completion rate
- Cost per successful task — the number that matters
How to roll out
A/B at light compression first (about 2×), then dial up if quality holds. Surface per-request receipts (tokens in/out), batch job reports, and dashboards so every stakeholder can see the trade-off.
Independent benchmark
FinanceBench.
FinanceBench, 128 questions over SEC filings. At light ~2× compression accuracy holds; push to ~10× when cost matters more than peak accuracy. Compression complements RAG — it is not a replacement.
Two ways to deploy
Pick the one that fits your stack.
Drop-in SDK. One API key.
Install, grab a key, compress any prompt or document before it hits your LLM. Pay per million tokens, no surprise bills.
- $10 in free credits on sign-up, no credit card required
- TypeScript & Python clients
- Question-aware compression
- Transparent per-million-token pricing
Sign up, get $10 of compression free, no card needed.
Runs inside your VPC.
Your data never leaves your network. We deploy Compresr to your infrastructure, tune it for your workload, and support you directly.
- Private deployment in your cloud or data center
- Custom throughput & latency SLAs
- Tailored to your business needs
- Dedicated support
Enterprise, finance, healthcare, regulated workloads.
FAQ
Questions, answered.
How can I reduce LLM token costs?
Start by identifying which context actually contributes to the answer and remove what does not. Compression, better retrieval, caching, shorter prompts, and model routing all compose — the biggest lever for most teams is cutting unnecessary input context before the model call.
What causes high token usage?
Large system prompts paid on every call, RAG over-retrieval, growing conversation history, verbose tool outputs, and full documents sent for narrow questions. Most of it is input context the model never needed.
How do input tokens affect LLM cost?
API cost scales with the number of tokens you send, and large prompts also add latency. Because input often dwarfs output, reducing input context is usually the highest-impact place to cut spend.
Can context compression reduce token costs?
Yes. Query-aware compression keeps the spans that carry the answer and drops the rest, often reducing input by roughly 20–60% with minimal quality loss depending on method and workload.
Does reducing tokens hurt accuracy?
It depends on the method. Blind truncation drops the tail where the answer often lives. Query-aware compression at light ratios can match or beat full-context accuracy — measure it with A/B before rollout.
How does prompt caching reduce token usage?
Caching stores repeated prefixes — stable system prompts and schemas — so they are not re-billed at full price on every call. It pairs well with compression on the variable part of the prompt.
How can I reduce RAG token costs?
Retrieve fewer, higher-signal chunks (de-dup, MMR, rerank) and compress what remains per query before it reaches the model, so you only pay for tokens that carry information.
How can I reduce AI agent token usage?
Compress verbose tool outputs, trim tool definitions, and reduce unnecessary steps that resubmit similar prefixes. A context gateway can auto-compress history and tool results inside the loop.
Cut your token bill without cutting quality.
Works with your existing stack. Start with light compression and measure — then dial it up when the numbers hold.
