Downsizing logodownsizing

APIs that make your agents better

Built for long-running agents. Route to the model that performs best, cut token consumption, catch agents when they underperform, and optimize inference costs end-to-end.

Empower your Claude agent

Works best with agents and AI coding tools

Takes 2 minutes or one prompt to set up.

Better performance, lower spend, for agents that run for hours. Use Anthropic / OpenAI by default — or route to alternative models at equivalent quality.

Talk to us →

Under the hood

We optimize at all levels

Inference API handles routing, compression, and model selection.

InputReduceFilterRouteOutput
KimiDeepSeekFableSonnetOpus
Input token blockReduced tokensFiltered tokensRouted output
01Input

Full token grid arrives - data, prompt and context.

02Reduce

We reduce input token volume by removing low-signal and redundant units.

03Filter

When feasible, we cherry-pick only the most salient pieces.

04Route

Smart router sends bundle to the best model the grader recommends.

05Output

Response compressed too.

Running a custom pipeline? Share your setup and we'll figure out what fits.

contact@downsizing.dev →

From engineers

Loved by our customers

We switched our agentic pipelines to Downsizing and started seeing savings right away.

Paula

Paula

Staff Engineer · Samsung

Went from $8k/month on Claude Code to $5k. Didn't change a single prompt.

Aleh

Aleh

Senior Software Engineer · Nvidia

Extends my Claude Code subscription at least by 2x on the free version. Love it.

Darshan

Darshan

Senior Software Engineer · Google

Downsizing replaced multiple plugins with better outcomes, 2 minutes to set up.

Anton

Anton

Senior Software Engineer · Revolut

Try it out. Install the CLI and connect coding tools or agents in seconds.

CLI quickstart →

Controls

Full control in real-time

Tune compression, routing, and attention for each workload — from a single config.

Input compression

Redundancy stripped from context before it reaches the model.

Output compression

Models instructed to be concise and eliminate scaffolding.

Smart routing

Match each request to the cheapest capable model.

Focused attention

Steers the model to attend only to the task-critical parts of context.

Adjust the controls above to configure your proxy — then sign up to save.

Analytics

Track savings in real-time

Actual Claude agentic sessions

downsizing.dev/dashboard/savings

Input tokens

3.5B

received

Input saved

1.44B

of 3.5B total

Requests

14,077

total

Output tokens

284M

received

Output saved

99.4M

of 284M total

Routed

8,602

61% of total

Input tokens by model

ModelRequestsTokens inSaved
Sonnet8,2412.1B
891M
Opus1,034412M
176M
Haiku4,802980M
372M

Recent requests

ModelSavedWhen
Sonnet+18,4322m ago
Haiku+4,2114m ago
Sonnet+34,2097m ago
Haiku+2,84711m ago

Benchmarked

Compression improves quality

In large contexts, it often improves it.

Claude Fable 5

Baseline

95%

Compressed

96.2%

Delta

+1.2%

Baseline
Compressed

Claude Opus 4.8

Baseline

88.6%

Compressed

89.9%

Delta

+1.3%

Baseline
Compressed

Claude Sonnet 5

Baseline

85.2%

Compressed

86.9%

Delta

+1.7%

Baseline
Compressed

Claude Haiku 4.5

Baseline

73.3%

Compressed

74.7%

Delta

+1.4%

Baseline
Compressed

GPT-5.6 Luna

Baseline

62.7%

Compressed

63.9%

Delta

+1.2%

Baseline
Compressed

GPT-5.6 Terra

Baseline

63.4%

Compressed

65.1%

Delta

+1.7%

Baseline
Compressed

GPT-5.6 Sol

Baseline

64.6%

Compressed

65.8%

Delta

+1.2%

Baseline
Compressed

*OpenAI has not published a SWE-bench Verified score for GPT-5.6. These baselines are OpenAI's official SWE-Bench Pro figures — a different, harder benchmark, not directly comparable to the SWE-bench Verified scores above.

With routing and compression enabled

The same as using Anthropic and OpenAI models at discounted prices

Processing mode
ModelInput / 1M tokOutput / 1M tok
Anthropic

Claude Fable 5

Mythos-class, most capable publicly available

$10$7.14
$50$38.46

Claude Opus 4.8

Most intelligent, for agents & coding

$5$3.57
$25$19.23

Claude Sonnet 5

Balance of intelligence, cost & speed

$3$2.14
$15$11.54

Claude Haiku 4.5

Fastest, most compact Anthropic model

$0.8$0.72
$4$3.60
OpenAI

GPT-5.6 Luna

Frontier model for complex reasoning

$5$3.57
$20$15.38

GPT-5.6 Terra

Affordable model for coding & work

$2.50$1.79
$15$11.54

GPT-5.6 Sol

Strongest mini for coding & sub-agents

$0.15$0.135
$0.6$0.54
Downsizing

Universe

Maximum capability, multimodal & deep reasoning

$1
$4

Galaxy

Balanced intelligence, optimized for agentic workloads

$0.5
$2

Star

Lightning-fast, ultra-low cost

$0.25
$1

*These figures represent typical compression gains observed across diverse applications and agents.