Large input context
Long system prompts, full documents, and oversized knowledge snippets are billed whenever they are sent.
Cut input tokens, not model quality. Keep the OpenAI models your application relies on while removing unnecessary context before each call.
Compresr sits before your OpenAI call. It shortens context; it does not replace OpenAI.
Add input optimization to your OpenAI stack
from compresr import Compresr from openai import OpenAI # Keep the task query separate. query = "What changed in operating margin?" # Compress only the long source context. result = compresr.compress( context=document, query=query ) # Send the shorter context to OpenAI. response = openai.responses.create( model=current_model, input=result.compressed_text )
Cost anatomy
OpenAI usage is priced per token, and every call adds input and output cost. In context-heavy applications, the same documents, histories, and tool results may be resent across many calls.
Per-task cost
Σ (input tokens × input rate + output tokens × output rate) across every model call
Cached input may receive discounted pricing, but uncached context and repeated workflow steps still accumulate. Measure the entire task, not an isolated request.
Long system prompts, full documents, and oversized knowledge snippets are billed whenever they are sent.
Shotgun retrieval sends duplicate or low-relevance chunks even when only a few spans support the answer.
Chat and agent workflows repeatedly submit prior turns, intermediate reasoning, and accumulated state.
Verbose API results, unnecessary tool calls, failures, and retries multiply tokens across a single task.
In RAG and agent systems, input tokens can dominate because large contexts are resent every turn. Start by mapping tokens by workflow stage, then remove context that does not change the answer.
Optimization stack
There is no single switch. The strongest production strategy combines context reduction, caching, retrieval tuning, sensible routing, and workflow measurement.
Use provider prompt caching for repeated system prompts and prefixes. Choose cache windows around actual reuse patterns.
Tighten filters, deduplicate results, reduce N, and rerank so only high-signal passages consume the prompt budget.
Keep guardrails explicit but concise, set useful response limits, and avoid feeding redundant streamed output back into the prompt.
Cap agent prompt size, prefer idempotent tools, compress tool responses, and eliminate failure loops that repeatedly spend tokens.
Send simple tasks to lower-cost models while preserving your current model for tasks where its quality is required.
Optimize net cost per successful task—not the apparent price of one call that may trigger retries or reduce completion rate.
How Compresr fits
Compresr takes a long context plus the task query and returns a shorter string. Your application then sends that reduced context to OpenAI using the same model and provider integration.
Original context
Compresr + query
Reduced context
OpenAI
Response
RAG chunks, long documents, search results, conversation history, agent tool outputs, and database or API results. Pass the user or task query so answer-bearing spans can be retained.
Integrate through REST, Python, or TypeScript. Single-call, streaming, and batch endpoints support interactive applications and high-throughput pipelines. Compresr does not call OpenAI for you.
View the OpenAI integration→Workflow playbooks
Reduce N, deduplicate and rerank chunks, then compress the remaining context against the user query before sending it to OpenAI.
Explore RAG cost optimization →AI agentsCompress verbose tool output and accumulated history, set per-turn prompt budgets, and prevent repeated payloads from expanding every loop.
Explore agent cost optimization →ChatbotsPeriodically compress long conversation history while retaining salient user facts, decisions, constraints, and unresolved requests.
Explore chatbot cost optimization →Document AIExtract answer-aligned passages from long PDFs, reports, and contracts instead of sending entire files for a targeted question.
Explore document AI optimization →Two different levers
A lower-cost OpenAI model can reduce spend, but it may also change output quality and behavior. Input optimization lets you test savings while keeping the model your application already trusts.
Change the model
Best when the task is simple enough for a lower-cost model. Requires evaluation because reasoning, instruction following, tool use, latency, and output style may change.
Keep the model
Keep your current OpenAI model and test whether shorter, query-relevant context lowers cost while preserving task quality. Re-evaluate model routing after input waste is removed.
Quality guardrails
Run held-out evaluations at multiple compression ratios. Plot cost against accuracy, completion rate, and latency to find the efficient frontier for each production workflow.
original_tokens
before
compressed_tokens
after
actual_compression_ratio
ratio
tokens_saved
delta
duration_ms
latency
Track OpenAI cost before and after, Compresr processing cost, call count, output tokens, and final net savings.
Evaluate accuracy on a held-out set, task completion, answer faithfulness, and the rate of retries or human escalation.
A/B light compression around 2× against stronger ratios. Select the lowest cost that still clears your quality threshold.
Independent benchmark
FinanceBench, 128 questions over SEC filings. At light ~2× compression accuracy holds; push to ~10× when cost matters more than peak accuracy. Compression complements RAG — it is not a replacement.
Two ways to deploy
Install, grab a key, compress any prompt or document before it hits your LLM. Pay per million tokens, no surprise bills.
Sign up, get $10 of compression free, no card needed.
Your data never leaves your network. We deploy Compresr to your infrastructure, tune it for your workload, and support you directly.
Enterprise, finance, healthcare, regulated workloads.
FAQ
Compress unnecessary input context, cache stable prefixes, tune retrieval, rerank results, cap outputs, shorten prompts, reduce agent steps, and route simple tasks only when quality permits. Measure cost per successful task.
Common causes include high volume, redundant input context, long histories, RAG over-retrieval, verbose tool output, multi-step loops, retries, and resending stable prefixes without caching.
OpenAI prices API use per one million input and output tokens, with discounted rates for eligible cached input. Rates vary by model, so verify the current provider pricing before forecasting.
Apply query-aware compression to long documents, RAG chunks, tool output, and histories. Tighten retrieval, remove duplicates, shorten verbose instructions, and avoid sending context unrelated to the current task.
Yes. Compresr returns a shorter string that your application places in the appropriate system or context slot. Keep the user question separate and continue calling OpenAI normally.
Yes. Start with context reduction and prompt caching on your current model. After validating quality, decide whether routing selected tasks to a cheaper model provides additional value.
Eligible cached input receives lower per-token rates. Cache stable, frequently reused prefixes and combine caching with compression to reduce the variable context that still changes between calls.
For RAG, retrieve less, rerank, and compress the selected text. For agents, compress tool outputs between steps, trim history, cap prompt budgets, and remove unnecessary calls or retries.
Keep OpenAI. Send less.
Benchmark your own OpenAI workload, compare tokens before and after, and validate quality before changing production traffic.