Double the subscription you already have
Extend your existing subscription with smart routing and compaction - or use open-source models at a discounted price.
Get started
Connect your subscription in minutes
No subscription? Pay as you go with API credits
You could route requests to open-source models - better prices and same quality for most tasks
Sign up →Coding Subscriptions
Bring your coding subscription, and we'll double its capacity
Claude, Codex, or any other agent
Bring your own
Freeunlimited usage
Bring your existing Claude, Codex, or other agent subscription and we'll extend your effective limits.
- Input & output token compression
- Smart routing within your provider
- Focused attention mode
- Usage dashboard & analytics
Our subscription
Full access — no external subscription required. We provide the tokens with doubled capacity compared to direct Anthropic or OpenAI plans.
- Everything in Bring your own
- 2× token capacity vs. Claude Code or Codex
No setup fees · Cancel anytime
Running Hermes or another long-lived agent? The longer it runs, the more compression saves you. Start with $10 in free credits.
Claim your credits →From engineers
Loved by our customers
“We switched our agentic pipelines to Downsizing and started seeing savings right away.”

Paula
Staff Engineer · Samsung
“Went from $8k/month on Claude Code to $5k. Didn't change a single prompt.”

Aleh
Senior Software Engineer · Nvidia
“Extends my Claude Code subscription at least by 2x on the free version. Love it.”

Darshan
Senior Software Engineer · Google
“Downsizing replaced multiple plugins with better outcomes, 2 minutes to set up.”

Anton
Senior Software Engineer · Revolut
Integrations
Works with SDKs, Claude Code, Codex and many others
No code changes, only configuration changes in your existing tools or workflows.
Coding & agentic tools
We silently compress redundant context so you can code longer and spend significantly less.
Personal AI Assistants
We prune stale context and use cheaper models for simple tasks so your agents run longer for less.
Claude Code
Official CLI for agentic coding
LiveCodex
OpenAI's terminal coding agent with custom endpoint support
LiveJunie
JetBrains AI coding agent for terminal and IDE
LiveHermes
Evolving AI agent with persistent memory
SoonAmp
Agentic coding tool from Sourcegraph with custom provider support
LiveCline
VS Code extension for code generation & file ops
LiveCrush
Terminal-based AI tool with CLI and TUI interfaces
LiveDroid
Enterprise terminal agent for end-to-end workflows
LiveEigent
Desktop multi-agent for browser automation
LiveGemini CLI
Google's terminal coding agent
LiveGoose
AI agent for local execution & engineering tasks
LiveOpenHands
Open-source autonomous coding agent platform
LivePi
Minimal terminal coding harness with unified LLM API
SoonWarp
AI-native terminal with Agent Mode and custom model routers
Using a tool not listed? Let us know — we can support it.
Same quality. Same models. Better prices.
See what your current plan gets you — with and without downsizing
At the very least, you should expect that much
| Plan | Price | Usage | With Downsizing |
|---|---|---|---|
| Anthropic | |||
| Pro | $17 / mo | 1× base usage | Up to 2× more usage |
| Max | From $100 / mo | 5× or 20× usage | Up to 10× or 40× usage |
| OpenAI | |||
| Plus | $20 / mo | Fixed message cap | Up to 2× more usage |
| Pro | $200 / mo | Unlimited (rate-limited) | Up to 2× more usage |
Controls
Full control in real-time
Tune compression, routing, and attention for each workload — from a single config.
Input compression
Redundancy stripped from context before it reaches the model.
Output compression
Models instructed to be concise and eliminate scaffolding.
Smart routing
Match each request to the cheapest capable model.
Focused attention
Steers the model to attend only to the task-critical parts of context.
Adjust the controls above to configure your proxy — then sign up to save.
Analytics
Track savings in real-time
Actual Claude agentic sessions
Input tokens
3.5B
received
Input saved
1.44B
of 3.5B total
Requests
14,077
total
Output tokens
284M
received
Output saved
99.4M
of 284M total
Routed
8,602
61% of total
Input tokens by model
| Model | Requests | Tokens in | Saved |
|---|---|---|---|
| Sonnet | 8,241 | 2.1B | 891M |
| Opus | 1,034 | 412M | 176M |
| Haiku | 4,802 | 980M | 372M |
Recent requests
| Model | Saved | When |
|---|---|---|
| Sonnet | +18,432 | 2m ago |
| Haiku | +4,211 | 4m ago |
| Sonnet | +34,209 | 7m ago |
| Haiku | +2,847 | 11m ago |
Benchmarked
Compression improves quality
In large contexts, it often improves it.
Claude Fable 5
Baseline
95%
Compressed
96.2%
Delta
+1.2%
Claude Opus 4.8
Baseline
88.6%
Compressed
89.9%
Delta
+1.3%
Claude Sonnet 5
Baseline
85.2%
Compressed
86.9%
Delta
+1.7%
Claude Haiku 4.5
Baseline
73.3%
Compressed
74.7%
Delta
+1.4%
GPT-5.6 Luna
Baseline
62.7%
Compressed
63.9%
Delta
+1.2%
GPT-5.6 Terra
Baseline
63.4%
Compressed
65.1%
Delta
+1.7%
GPT-5.6 Sol
Baseline
64.6%
Compressed
65.8%
Delta
+1.2%
*OpenAI has not published a SWE-bench Verified score for GPT-5.6. These baselines are OpenAI's official SWE-Bench Pro figures — a different, harder benchmark, not directly comparable to the SWE-bench Verified scores above.