Guide

How to Make Claude Code Cheaper: 7 Ways to Cut Your Bill

7 min read

Claude Code is a superb coding agent, and it can get expensive fast. Most of the cost comes from one thing: tokens. Every message re-sends your conversation and the files in context, and the most capable model bills the most per token. The good news is that most of the spend is controllable. Here are seven ways to cut your Claude Code bill, from quick wins to bigger decisions, with honest tradeoffs for each.

Why does Claude Code get expensive?

Claude Code charges for the tokens it reads (your prompt, conversation history, and any files in context) and the tokens it writes. Two things push the bill up: using the most powerful model for routine work, and letting the context grow so you pay to re-read it on every turn. Almost every tactic below targets one of those two levers.

1. Use the right model for the task

Claude Code can run on different models, and they are not priced the same. Opus is the most capable and the most expensive. Sonnet is the balance of speed and intelligence at a lower price. Haiku is the fastest and cheapest. You do not need Opus to rename variables, write a test, or fix a small bug.

Switch models with the /model command inside Claude Code: drop to Sonnet or Haiku for routine edits, and reserve Opus for genuinely hard, multi-file reasoning. This single habit often cuts a bill more than anything else.

2. Keep your context small

The bigger your conversation and the more files loaded into context, the more tokens Claude Code re-reads on every message. Long sessions quietly get expensive. Two built-in commands help:

  • /clear starts a fresh conversation when you switch to an unrelated task, so you stop paying to carry old history.
  • /compact summarizes the conversation so far, shrinking the tokens re-sent each turn while keeping the gist.

Also point Claude Code at the specific files that matter instead of dumping the whole repository into context.

3. Let prompt caching work for you

Claude Code caches stable context so it does not pay full price to re-read the same thing every turn. Cached tokens cost a small fraction of fresh input. You benefit from this automatically, but you can help it: keep a task in one continuous session rather than restarting constantly, and avoid churning the files in context, since changing the earlier part of a prompt invalidates the cache after that point.

4. Pick the right billing: subscription or API

This is the biggest decision for many people. Claude Code runs two ways:

  • A Claude subscription (Pro or Max): a flat monthly fee with generous usage. If you use Claude Code most days, a subscription is usually far cheaper than paying per token.
  • An API key: pay-as-you-go per token. Best if you use it occasionally or in bursts, where a monthly subscription would sit idle.

Match the plan to how you actually work. Heavy daily user paying per token? Move to a subscription. Light or spiky usage on a subscription you rarely hit? The API may be cheaper.

5. Plan the task before you prompt

Long, vague back-and-forth is expensive because each turn re-sends a growing context. State the task, the intent, and the constraints clearly up front, and let the agent run. A well-specified single request often costs less and works better than ten corrective follow-ups, each paying to re-read the whole thread.

6. Turn down effort on easy work

More reasoning means more tokens. For simple, well-scoped tasks you do not need the model thinking at maximum depth. Save the heavy settings for genuinely hard problems, and let routine work run lean.

7. Route through a cheaper compatible proxy

If you have optimized the above and the bill is still high, you can point Claude Code at a compatible proxy that routes your requests to a cheaper model. That is what KairoTokens does: you set two environment variables and keep your exact Claude Code workflow, but requests run through a cost-efficient frontier model at a fraction of the price. You buy prepaid credits, there is no subscription, and delivery is instant.

The honest tradeoff: a proxy is not Anthropic and does not run Anthropic's Claude models. KairoTokens is an independent service, not affiliated with Anthropic. You are trading the exact Claude model for a compatible, cheaper one that keeps the same tools and workflow. For a lot of everyday coding, that tradeoff is worth it; for work where you specifically want Claude's model, keep using it directly and lean on tactics 1 through 6. See the setup guide for how to point Claude Code at it, and the payment guide if you are new to crypto.

Is a Claude Code proxy safe to use?

A proxy sits between Claude Code and the model, so it can see the requests you send. Before trusting any proxy, check a few things: does it clearly state who runs it and what it does with your data, can you verify it is working (KairoTokens has a public balance checker so you can confirm your key and usage), and is the billing transparent. Treat your proxy key like any API key: keep it private, and only load funds you are comfortable spending. If a service hides what it is or makes claims that seem too good to be true, walk away.


The short version

  • Right model: use /model to drop to Sonnet or Haiku for routine work; save Opus for hard problems.
  • Small context: use /clear and /compact; load only the files you need.
  • Caching: keep tasks in one session so stable context stays cheap.
  • Billing: heavy daily use, get a subscription; light or spiky use, pay per token via API.
  • Plan first: one clear, complete request beats ten corrective ones.
  • Lower effort on easy tasks.
  • Compatible proxy: route Claude Code through a cheaper model with a service like KairoTokens if you accept it is not Anthropic's own model.