Downsizing logodownsizing

Double the subscription you already have

Extend your existing subscription with smart routing and compaction - or use open-source models at a discounted price.

Works with
Claude CodeCursorCodexCopilotJunieHermesClineCrushDroidEigentGemini CLIGooseOpenHandsContinueKilo CodeOpenCodeRoo CodeZedn8nDifyClaude CodeCursorCodexCopilotJunieHermesClineCrushDroidEigentGemini CLIGooseOpenHandsContinueKilo CodeOpenCodeRoo CodeZedn8nDify

Get started

Connect your subscription in minutes

Get started free →$10 in free inference credits on signup

No subscription? Pay as you go with API credits

You could route requests to open-source models - better prices and same quality for most tasks

Sign up →

Coding Subscriptions

Bring your coding subscription, and we'll double its capacity

Claude, Codex, or any other agent

Bring your own

Freeunlimited usage

Bring your existing Claude, Codex, or other agent subscription and we'll extend your effective limits.

  • Input & output token compression
  • Smart routing within your provider
  • Focused attention mode
  • Usage dashboard & analytics
Get started free
Coming soon

Our subscription

Pro$15/mo
Max$70/mo

Full access — no external subscription required. We provide the tokens with doubled capacity compared to direct Anthropic or OpenAI plans.

  • Everything in Bring your own
  • 2× token capacity vs. Claude Code or Codex
Join waitlist

No setup fees · Cancel anytime

Running Hermes or another long-lived agent? The longer it runs, the more compression saves you. Start with $10 in free credits.

Claim your credits →

From engineers

Loved by our customers

We switched our agentic pipelines to Downsizing and started seeing savings right away.

Paula

Paula

Staff Engineer · Samsung

Went from $8k/month on Claude Code to $5k. Didn't change a single prompt.

Aleh

Aleh

Senior Software Engineer · Nvidia

Extends my Claude Code subscription at least by 2x on the free version. Love it.

Darshan

Darshan

Senior Software Engineer · Google

Downsizing replaced multiple plugins with better outcomes, 2 minutes to set up.

Anton

Anton

Senior Software Engineer · Revolut

Same quality. Same models. Better prices.

See what your current plan gets you — with and without downsizing

At the very least, you should expect that much

PlanPriceUsageWith Downsizing
Anthropic
Pro$17 / mo1× base usageUp to 2× more usage
MaxFrom $100 / mo5× or 20× usageUp to 10× or 40× usage
OpenAI
Plus$20 / moFixed message capUp to 2× more usage
Pro$200 / moUnlimited (rate-limited)Up to 2× more usage

Controls

Full control in real-time

Tune compression, routing, and attention for each workload — from a single config.

Input compression

Redundancy stripped from context before it reaches the model.

Output compression

Models instructed to be concise and eliminate scaffolding.

Smart routing

Match each request to the cheapest capable model.

Focused attention

Steers the model to attend only to the task-critical parts of context.

Adjust the controls above to configure your proxy — then sign up to save.

Analytics

Track savings in real-time

Actual Claude agentic sessions

downsizing.dev/dashboard/savings

Input tokens

3.5B

received

Input saved

1.44B

of 3.5B total

Requests

14,077

total

Output tokens

284M

received

Output saved

99.4M

of 284M total

Routed

8,602

61% of total

Input tokens by model

ModelRequestsTokens inSaved
Sonnet8,2412.1B
891M
Opus1,034412M
176M
Haiku4,802980M
372M

Recent requests

ModelSavedWhen
Sonnet+18,4322m ago
Haiku+4,2114m ago
Sonnet+34,2097m ago
Haiku+2,84711m ago

Benchmarked

Compression improves quality

In large contexts, it often improves it.

Claude Fable 5

Baseline

95%

Compressed

96.2%

Delta

+1.2%

Baseline
Compressed

Claude Opus 4.8

Baseline

88.6%

Compressed

89.9%

Delta

+1.3%

Baseline
Compressed

Claude Sonnet 5

Baseline

85.2%

Compressed

86.9%

Delta

+1.7%

Baseline
Compressed

Claude Haiku 4.5

Baseline

73.3%

Compressed

74.7%

Delta

+1.4%

Baseline
Compressed

GPT-5.6 Luna

Baseline

62.7%

Compressed

63.9%

Delta

+1.2%

Baseline
Compressed

GPT-5.6 Terra

Baseline

63.4%

Compressed

65.1%

Delta

+1.7%

Baseline
Compressed

GPT-5.6 Sol

Baseline

64.6%

Compressed

65.8%

Delta

+1.2%

Baseline
Compressed

*OpenAI has not published a SWE-bench Verified score for GPT-5.6. These baselines are OpenAI's official SWE-Bench Pro figures — a different, harder benchmark, not directly comparable to the SWE-bench Verified scores above.