YBacked by Y Combinator

How to Reduce OpenAI API Costs in Production

Cut input tokens, not model quality. Keep the OpenAI models your application relies on while removing unnecessary context before each call.

  • Reduce OpenAI input tokens across RAG, agents, chatbots, and documents.
  • Keep your current OpenAI models and existing application architecture.
  • Benchmark cost, latency, and quality before production rollout.
PyTS$ pip install compresr

Compresr sits before your OpenAI call. It shortens context; it does not replace OpenAI.

Add input optimization to your OpenAI stack

_ OPENAI INTEGRATIONCopy prompt

from compresr import Compresr from openai import OpenAI # Keep the task query separate. query = "What changed in operating margin?" # Compress only the long source context. result = compresr.compress( context=document, query=query ) # Send the shorter context to OpenAI. response = openai.responses.create( model=current_model, input=result.compressed_text )

Cost anatomy

Why OpenAI API costs increase.

OpenAI usage is priced per token, and every call adds input and output cost. In context-heavy applications, the same documents, histories, and tool results may be resent across many calls.

Per-task cost

Σ (input tokens × input rate + output tokens × output rate) across every model call

Cached input may receive discounted pricing, but uncached context and repeated workflow steps still accumulate. Measure the entire task, not an isolated request.

Large input context

Long system prompts, full documents, and oversized knowledge snippets are billed whenever they are sent.

RAG over-retrieval

Shotgun retrieval sends duplicate or low-relevance chunks even when only a few spans support the answer.

Growing history

Chat and agent workflows repeatedly submit prior turns, intermediate reasoning, and accumulated state.

Tool and retry loops

Verbose API results, unnecessary tool calls, failures, and retries multiply tokens across a single task.

In RAG and agent systems, input tokens can dominate because large contexts are resent every turn. Start by mapping tokens by workflow stage, then remove context that does not change the answer.

Optimization stack

Ways to reduce OpenAI API costs.

There is no single switch. The strongest production strategy combines context reduction, caching, retrieval tuning, sensible routing, and workflow measurement.

Cache stable prefixes

Use provider prompt caching for repeated system prompts and prefixes. Choose cache windows around actual reuse patterns.

Retrieve less, then rerank

Tighten filters, deduplicate results, reduce N, and rerank so only high-signal passages consume the prompt budget.

Constrain prompts and output

Keep guardrails explicit but concise, set useful response limits, and avoid feeding redundant streamed output back into the prompt.

Remove unnecessary steps

Cap agent prompt size, prefer idempotent tools, compress tool responses, and eliminate failure loops that repeatedly spend tokens.

Route only when quality allows

Send simple tasks to lower-cost models while preserving your current model for tasks where its quality is required.

Measure successful tasks

Optimize net cost per successful task—not the apparent price of one call that may trigger retries or reduce completion rate.

How Compresr fits

We keep the signal and drop the noise.

Compresr takes a long context plus the task query and returns a shorter string. Your application then sends that reduced context to OpenAI using the same model and provider integration.

01

Original context

→
02

Compresr + query

→
03

Reduced context

→
04

OpenAI

→
05

Response

What goes through Compresr

RAG chunks, long documents, search results, conversation history, agent tool outputs, and database or API results. Pass the user or task query so answer-bearing spans can be retained.

One layer in your stack

Integrate through REST, Python, or TypeScript. Single-call, streaming, and batch endpoints support interactive applications and high-throughput pipelines. Compresr does not call OpenAI for you.

View the OpenAI integration→

Two different levers

Compression before model switching.

A lower-cost OpenAI model can reduce spend, but it may also change output quality and behavior. Input optimization lets you test savings while keeping the model your application already trusts.

Change the model

Route to a cheaper model

Best when the task is simple enough for a lower-cost model. Requires evaluation because reasoning, instruction following, tool use, latency, and output style may change.

Keep the model

Reduce unnecessary input first

Keep your current OpenAI model and test whether shorter, query-relevant context lowers cost while preserving task quality. Re-evaluate model routing after input waste is removed.

Compare context compression and cheaper models →

Quality guardrails

Measure savings against quality.

Run held-out evaluations at multiple compression ratios. Plot cost against accuracy, completion rate, and latency to find the efficient frontier for each production workflow.

PER-REQUEST RECEIPTmeasured

original_tokens

before

compressed_tokens

after

actual_compression_ratio

ratio

tokens_saved

delta

duration_ms

latency

Net cost

Track OpenAI cost before and after, Compresr processing cost, call count, output tokens, and final net savings.

Quality

Evaluate accuracy on a held-out set, task completion, answer faithfulness, and the rate of retries or human escalation.

Efficient frontier

A/B light compression around 2× against stronger ratios. Select the lowest cost that still clears your quality threshold.

Independent benchmark

FinanceBench.

BaselineGPT-5.2
latte_v2 API+ GPT-5.2
Compression
~2×
Average context
~106K
~56K
Accuracy
73%
77%
Savings
~47% cheaper

FinanceBench, 128 questions over SEC filings. At light ~2× compression accuracy holds; push to ~10× when cost matters more than peak accuracy. Compression complements RAG — it is not a replacement.

Two ways to deploy

Pick the one that fits your stack.

Hosted SDK

Drop-in SDK. One API key.

Install, grab a key, compress any prompt or document before it hits your LLM. Pay per million tokens, no surprise bills.

  • $10 in free credits on sign-up, no credit card required
  • TypeScript & Python clients
  • Question-aware compression
  • Transparent per-million-token pricing
Get your free credits

Sign up, get $10 of compression free, no card needed.

On-prem deployment

Runs inside your VPC.

Your data never leaves your network. We deploy Compresr to your infrastructure, tune it for your workload, and support you directly.

  • Private deployment in your cloud or data center
  • Custom throughput & latency SLAs
  • Tailored to your business needs
  • Dedicated support
Contact us for on-prem

Enterprise, finance, healthcare, regulated workloads.

FAQ

OpenAI cost questions, answered.

How can I reduce OpenAI API costs?

Compress unnecessary input context, cache stable prefixes, tune retrieval, rerank results, cap outputs, shorten prompts, reduce agent steps, and route simple tasks only when quality permits. Measure cost per successful task.

What causes high OpenAI API bills?

Common causes include high volume, redundant input context, long histories, RAG over-retrieval, verbose tool output, multi-step loops, retries, and resending stable prefixes without caching.

How do input tokens affect OpenAI pricing?

OpenAI prices API use per one million input and output tokens, with discounted rates for eligible cached input. Rates vary by model, so verify the current provider pricing before forecasting.

How can I reduce OpenAI input tokens?

Apply query-aware compression to long documents, RAG chunks, tool output, and histories. Tighten retrieval, remove duplicates, shorten verbose instructions, and avoid sending context unrelated to the current task.

Does context compression work with OpenAI?

Yes. Compresr returns a shorter string that your application places in the appropriate system or context slot. Keep the user question separate and continue calling OpenAI normally.

Can I reduce costs without switching OpenAI models?

Yes. Start with context reduction and prompt caching on your current model. After validating quality, decide whether routing selected tasks to a cheaper model provides additional value.

How does prompt caching affect OpenAI costs?

Eligible cached input receives lower per-token rates. Cache stable, frequently reused prefixes and combine caching with compression to reduce the variable context that still changes between calls.

How do I reduce OpenAI costs for RAG and agents?

For RAG, retrieve less, rerank, and compress the selected text. For agents, compress tool outputs between steps, trim history, cap prompt budgets, and remove unnecessary calls or retries.

Keep OpenAI. Send less.

Stop paying for context your model doesn’t need.

Benchmark your own OpenAI workload, compare tokens before and after, and validate quality before changing production traffic.