YBacked by Y Combinator

Reduce LLM token costs in production AI applications.

Reduce the tokens your app sends without removing the information needed to complete the task. Compresr drops unnecessary input context before the model call — so you cut cost and latency across RAG, agents, chat, and docs.

  • Works with your existing LLM stack — OpenAI, Anthropic, Gemini.
  • Cuts input-token waste across RAG, agents, chat, and docs.
  • Measure savings vs. quality before rollout.
PyTS$ pip install compresr

Works with OpenAI, Anthropic, Gemini — provider-agnostic.

See it on your own file in 60 seconds

PASTE INTO CLAUDE CODECopy prompt

Use compresr to show live cost savings on my file. 1. pip install compresr 2. Introspect the SDK to discover the API. 3. Ask me for COMPRESR_API_KEY (open https://compresr.ai/signup: $10 free, no card). 4. Ask me for (a) a long document (PDF/.md/.txt) and (b) a question about it. 5. Compress with the question, then print a receipt: tokens in/out, ratio, % saved. 6. Ask the model the question against the compressed context and print the answer.

Why it adds up

Why LLM token costs get expensive.

Long context windows encourage larger prompts, but cost scales with tokens and adds latency. Most of that spend is input context the model never needed.

Long context windows

Bigger windows invite bigger prompts — but you pay for every token you send, whether the model reads it or not.

Large system prompts

Static instructions and schemas are paid on every single call, multiplying across your entire traffic volume.

RAG over-retrieval

Retrieving too many chunks “just in case” sends far more context to the model than the answer requires.

Conversation history

Accumulating multi-turn history bloats every request, resubmitting the same tokens turn after turn.

Verbose tool outputs

Agent loops re-enter raw tool results — search hits, API dumps, logs — that are mostly irrelevant to the next step.

Full documents for narrow questions

A whole PDF is sent to answer one line item, and multi-step agents resubmit similar prefixes repeatedly.

Where it hides

Where token waste happens.

Workload
Common token waste
RAG
Too many retrieved chunks
Agents
Large, verbose tool outputs
Chatbots
Growing conversation history
Documents
Full files sent for narrow questions
Research
Search results and scraped content

The toolkit

Ways to reduce LLM token costs.

These methods compose. Compression pairs well with caching, retrieval trimming, and model routing — you rarely pick just one.

01

Context compression

Query-aware compression keeps the spans that carry the answer and drops the rest.

02

Better retrieval

De-dup, Max Marginal Relevance, and smaller chunk budgets stop over-retrieval at the source.

03

Reranking

Rerank before model input so only the highest-signal passages ever reach the prompt.

04

Prompt caching

Cache repeated prefixes so stable instructions are not re-billed on every call.

05

Shorter system prompts

Slim instructions and tool schemas trim tokens paid on every request.

06

Model routing

Route to smaller, cheaper models when the task allows without losing quality.

07

Output limits

Stop sequences and max-token caps keep generation from drifting past what you need.

08

Fewer agent steps

Reduce unnecessary tool calls and loop iterations that resubmit similar context.

How it works

We keep the signal and drop the noise.

Compresr sits in front of your existing stack. Feed it the query and the context; it returns a shorter context you forward to your LLM — nothing else changes.

Original context112,552tokens in
Compresr226×query-aware
Reduced context498tokens out
Existing LLMAnyprovider

What Compresr can compress

  • RAG chunks before they reach the model
  • Long documents, per query
  • Growing conversation history
  • Search results and scraped content
  • Tool outputs and database results

The principle

The goal is not maximum compression. It is the lowest token usage that still meets the application's quality requirements. Drop in via SDK or API, keep OpenAI, Anthropic, or Gemini, and use batch endpoints for pipelines — streaming available.

What most teams are losing

Stop overpaying.

If you are paying full price for your tokens, you are leaving real money on the table. Measure savings against quality before you roll out.

~90%

Bill cut

vs. sending the full context

10×

Avg. compression

on long, sparse context

+3.7pp

Accuracy uplift

Compresr 2× vs full context, Pax arena

What to track

  • Tokens before vs. after, and net model cost
  • Compression cost and resulting net savings
  • Accuracy and task-completion rate
  • Cost per successful task — the number that matters

How to roll out

A/B at light compression first (about 2×), then dial up if quality holds. Surface per-request receipts (tokens in/out), batch job reports, and dashboards so every stakeholder can see the trade-off.

Independent benchmark

FinanceBench.

BaselineGPT-5.2
latte_v2 API+ GPT-5.2
Compression
None
~2×
Average context
~106K
~56K
Accuracy
73%
77%
Savings
None
~47% cheaper

FinanceBench, 128 questions over SEC filings. At light ~2× compression accuracy holds; push to ~10× when cost matters more than peak accuracy. Compression complements RAG — it is not a replacement.

Two ways to deploy

Pick the one that fits your stack.

Hosted SDK

Drop-in SDK. One API key.

Install, grab a key, compress any prompt or document before it hits your LLM. Pay per million tokens, no surprise bills.

  • $10 in free credits on sign-up, no credit card required
  • TypeScript & Python clients
  • Question-aware compression
  • Transparent per-million-token pricing
Get your free credits

Sign up, get $10 of compression free, no card needed.

On-prem deployment

Runs inside your VPC.

Your data never leaves your network. We deploy Compresr to your infrastructure, tune it for your workload, and support you directly.

  • Private deployment in your cloud or data center
  • Custom throughput & latency SLAs
  • Tailored to your business needs
  • Dedicated support
Contact us for on-prem

Enterprise, finance, healthcare, regulated workloads.

FAQ

Questions, answered.

How can I reduce LLM token costs?

Start by identifying which context actually contributes to the answer and remove what does not. Compression, better retrieval, caching, shorter prompts, and model routing all compose — the biggest lever for most teams is cutting unnecessary input context before the model call.

What causes high token usage?

Large system prompts paid on every call, RAG over-retrieval, growing conversation history, verbose tool outputs, and full documents sent for narrow questions. Most of it is input context the model never needed.

How do input tokens affect LLM cost?

API cost scales with the number of tokens you send, and large prompts also add latency. Because input often dwarfs output, reducing input context is usually the highest-impact place to cut spend.

Can context compression reduce token costs?

Yes. Query-aware compression keeps the spans that carry the answer and drops the rest, often reducing input by roughly 20–60% with minimal quality loss depending on method and workload.

Does reducing tokens hurt accuracy?

It depends on the method. Blind truncation drops the tail where the answer often lives. Query-aware compression at light ratios can match or beat full-context accuracy — measure it with A/B before rollout.

How does prompt caching reduce token usage?

Caching stores repeated prefixes — stable system prompts and schemas — so they are not re-billed at full price on every call. It pairs well with compression on the variable part of the prompt.

How can I reduce RAG token costs?

Retrieve fewer, higher-signal chunks (de-dup, MMR, rerank) and compress what remains per query before it reaches the model, so you only pay for tokens that carry information.

How can I reduce AI agent token usage?

Compress verbose tool outputs, trim tool definitions, and reduce unnecessary steps that resubmit similar prefixes. A context gateway can auto-compress history and tool results inside the loop.

Cut your token bill without cutting quality.

Works with your existing stack. Start with light compression and measure — then dial it up when the numbers hold.