← All guides
TroubleshootingToken usage

Claude Code Slow? Why It Happens and How to Speed It Up

Neo ZinoBy Neo Zino - builder of ClockedCode12 min read

Claude Code feels slow for five different reasons: a cold cache, bloated context, high effort, a heavy setup, or memory. Here is how to tell which and fix it.

Claude Code Slow? Why It Happens and How to Speed It Up

Made with DispatchSEO

On this page

Claude Code is slow for one of five reasons: the prompt cache went cold, the context has grown huge, the effort level makes it think longer than the task needs, something in your setup (an MCP server, hook or plugin) adds work, or the process is short on memory. You can tell which in about two minutes with /usage, /context and claude --safe-mode, before you change a single setting.

TL;DR: Slow on the first reply after a break means a cold cache. Slow only in long sessions means context, so /clear or /compact. A long pause before any text means thinking time, so /effort low. Slow from the first prompt of every session means a customization, so bisect with claude --safe-mode. Fans spinning means memory, and Claude Code warns you past 2.5GB. If --safe-mode is just as slow and the session is small, it is probably not you.

What you see, what it usually is, what to run

  • First reply after a break is slow, then it speeds up

    Cold prompt cache: the whole context was reprocessed

    /usage - read the Prompt cache line

  • Fine at 9am, crawling by 3pm in the same session

    Context has grown; every turn re-sends all of it

    /context, then /compact or /clear

  • Long pause before any text appears

    Thinking time from a high effort level or a heavy model

    /effort and /model

  • Slow from the very first prompt in every session

    An MCP server, hook or plugin is adding work

    claude --safe-mode

  • Fans spin, terminal lags, memory warning appears

    Process memory passed 2.5GB

    /compact, or /heapdump

Source: code.claude.com/docs/en/troubleshooting, /costs and /model-config, checked 2026-09-29.

Is it slow everywhere, or only in this one session?

Open a brand new session in the same project and send a one-line prompt. That single test splits the problem in half.

If the new session is fast, your setup is fine and the old session is the problem. That means context size or a cold cache, and the fix is on the session, not on your config. If the new session is slow too, run claude --safe-mode. Per the CLI help it starts with all customizations off: CLAUDE.md, skills, installed plugins, hooks, MCP servers, custom commands and agents, and output styles. Fast in safe mode means one of your customizations is the cause. Slow in safe mode as well means look at effort level, model, memory, or the service itself.

Why the first message after a break crawls

This is the cause I rarely see named in the threads on this topic, and it explains the most confusing version of the complaint: "it was fine, I went to lunch, now it's awful, and two messages later it's fine again."

Claude Code sends your full conversation with every request and relies on prompt caching so it does not reprocess all of it each time. The cache has a lifetime. Anthropic's costs page says your first message after a break longer than that lifetime misses the cache and reprocesses your full context. On a long session that is a big request, and it feels like the tool froze.

How long you can step away before the next reply is slow

Subscription (Pro, Max, Team, Enterprise)

1 hour

Subscription drawing on usage credits

5 minutes

API key or cloud provider (default)

5 minutes

Come back after the window and your first message reprocesses the full context instead of reading it from cache.

Source: code.claude.com/docs/en/costs, checked 2026-09-29.

You can check whether this is happening to you. From version 2.1.251, /usage adds a Prompt cache (main) line to the session block. The docs show it looking like this:

Prompt cache (main):   14 requests · 91% of input tokens from cache · 2 misses (last 6m 10s ago, 310.2k tokens re-cached) · 1 expected rebuild (compaction or tool-result clearing) · warm (1h TTL, last activity 40s ago)

I did not generate that output myself, because this run has no logged-in session to measure against. It is the format from Anthropic's docs. The parts to read are the misses count with its "last N ago" time and the warm or cold flag. If the last miss lines up with when you came back from a break, you have your answer. From version 2.1.260 the line can also name a likely cause, such as tool definitions changing.

What to do about it:

  • If you are about to step away from a huge session for longer than the window, that is a good moment to /compact or /clear instead of leaving it for later.
  • On Pro and Max, when you resume a large session after a long break, Claude Code offers to resume from a summary, so later requests do not carry the full history. Take that offer when you do not need the exact transcript.
  • If you draw on usage credits and want the one-hour lifetime back, the docs describe choosing the TTL yourself.

Context that keeps growing

Every turn re-sends the whole conversation, including every file Claude read and every tool result it kept. A session that took two seconds per reply at the start does not stay there. Our context window guide covers what fills it, and how auto-compact works covers what happens when it gets close to full.

The practical loop is short:

  1. Run /context to see what is using space.
  2. If you are switching to unrelated work, /rename the session and /clear. Anthropic's costs page says stale context wastes tokens on every subsequent message, and you can come back with /resume.
  3. If it is the same task, /compact with a focus, for example /compact keep only the plan and the diff.
  4. Ask Claude to read big files in a line range or a single function, and push noisy work like running a test suite into a subagent so the output stays out of your main context.

If compaction itself is the thing that hangs, that is a different problem with its own fixes for Claude Code stuck on compacting.

A permanent status line makes this much easier to catch early. The Claude Code status line builder can show context usage all the time, so you notice a session getting heavy before it feels slow.

Thinking time: effort level, model, and fast mode

A long pause before the first word usually is not the network. It is thinking. Extended thinking is on by default, and thinking tokens are generated before the visible answer, so a high effort level on a simple task is pure latency.

Set effort with /effort (or the slider in /model). The levels differ by model, but the shape is the same: low for quick exchanges where you review each result, medium for day-to-day work with a clear scope, high where verification matters, up to max for hard problems you want worked through unattended. If you never touch it, you are on the model's default, which the docs list as medium on Opus 5.5 and Sonnet 5.5 and high on most other models.

/effort low
/model sonnet

Two things worth knowing before you go hunting for a switch:

  • You cannot turn thinking off on Opus 5.5, Sonnet 5.5 or the Fable models. On those, effort level is the lever, not a thinking toggle.
  • /model sonnet or haiku for simple work is the other lever. Anthropic's own guidance is to reserve Opus for complex architectural decisions and multi-step reasoning.

If you want Opus and want it quicker, fast mode is the paid answer. /fast toggles it. Anthropic describes it as the same Opus quality with a different API configuration, up to 2.5x faster, at a higher price: $8 input and $40 output per million tokens on Opus 5.5. On subscriptions it is billed from usage credits, not your plan's included usage, so it is a real cost decision. It also composes with a lower effort level for maximum speed on straightforward tasks. It is a research preview, so the details can change.

Speed fixes in the order I would try them

  1. 1Start clean between unrelated tasks

    /rename then /clear

    Free. Old work stays reachable with /resume.

  2. 2Lower the effort level for simple work

    /effort low

    Less thinking, so weaker on hard problems.

  3. 3Bisect your customizations

    claude --safe-mode

    Session runs without CLAUDE.md, hooks, MCP and plugins.

  4. 4Pay for speed on Opus

    /fast

    Up to 2.5x faster, higher per-token price, usage credits only on subscriptions.

Source: code.claude.com/docs/en/costs, /model-config and /fast-mode, checked 2026-09-29.

Your own setup: bisecting MCP servers, hooks and plugins

If claude --safe-mode is fast and normal claude is slow, one customization is the cause, and the job is to find which one. The docs point at Debug your configuration for the clean-configuration test. My order of suspects:

  1. MCP servers. Run /mcp and disable any you are not actively using. Tool definitions are deferred by default now, so only names and server instructions enter context up front, but a server that starts slowly or answers slowly still costs you time. There is more in our guide to keeping MCP servers from eating context. Anthropic also notes that CLI tools like gh and aws are more context-efficient than an equivalent MCP server because they add no per-tool listing.
  2. Hooks. A hook that runs on every tool call adds its own runtime every time. Time the script directly in your shell.
  3. Plugins. Same logic: disable one at a time.
  4. CLAUDE.md. Anthropic suggests keeping it under 200 lines and moving workflow-specific instructions into skills, which load only when invoked. It is the cheapest trim, but rarely the main culprit.

To see what is actually happening, start with claude --debug-file ./claude-debug.txt and read the log. It is the same flag the docs use to verify hooks.

High CPU and memory: the 2.5GB warning

When the terminal itself lags, fans spin and typing gets sluggish, it is memory. Claude Code shows a critical memory usage warning when a session's heap passes 2.5GB. The documented recovery is to restart Claude Code and run claude --continue, which resumes the conversation in a fresh process. /compact also frees memory outside fullscreen rendering. Restarting between major tasks and keeping large build directories in .gitignore are on the same list.

If memory stays high afterward, /heapdump writes two files to ~/Desktop: a heap snapshot and a -diagnostics.json memory breakdown. The command is hidden from the command menu, so type it in full. Attach only the diagnostics file to a public issue. The docs are explicit that the .heapsnapshot contains your full conversation and credentials.

One more oddity: a Markdown table over 200 rows renders only its first 200 rows in the terminal. Before version 2.1.208 it rendered every row, so resuming a session that held a huge table could stall while re-rendering.

Slow search, especially on WSL

If the slowness is specifically in searching or @file mentions, two documented causes apply.

On WSL, reading across the Windows and Linux filesystems has a disk performance penalty, and search can return fewer matches than expected. The fix is to keep your project on the Linux filesystem under /home/ instead of /mnt/c/, or make searches more specific ("find md5 use in JS files"). Note that claude doctor still shows Search as OK in that case. Our Claude Code on WSL guide goes through the rest of the WSL gotchas.

If the bundled ripgrep will not run on your system, install your platform's package and set USE_BUILTIN_RIPGREP to 0 in your environment or in the env block of settings.json. Then run claude doctor and check that the Search line shows your system ripgrep instead of OK (bundled).

When the slowness is on Anthropic's side

Sometimes it really is not you. The signs: claude --safe-mode in a fresh session is just as slow, /context shows a small session, and you see retries or an API Error: 529 Overloaded. Our 529 overloaded guide covers that error. In that state no local setting helps, and the useful move is to check Anthropic's status page and try again later.

If the slowdown looks like your usage draining rather than replies lagging, that is a different symptom with its own page: Claude Code hitting usage limits too fast.

When none of this applies

I have not benchmarked these fixes against each other, so treat the order in the ladder as my judgment, not a measurement. Speed also varies by model, load and prompt, and no doc gives a guaranteed number for any of these changes except fast mode's "up to 2.5x". A few limits worth being straight about:

  • If your prompt asks Claude to read half the repo, it will be slow on any settings. A vaguer request means broader scanning, so a specific one is faster.
  • Lowering effort trades speed for quality on hard problems. Do not run low on a tricky bug fix and then blame the model.
  • Fast mode costs real money on a subscription. Do not leave it on for long autonomous runs where speed does not matter.
  • If none of the checks change anything, run /doctor, then /feedback from inside the slow session to send it to Anthropic.

FAQ

Why is Claude Code so slow all of a sudden?

Usually one of four things changed: the conversation grew and every turn now re-sends all of it, you came back after a break longer than the prompt cache lifetime so the first reply reprocessed everything, the effort level or model is doing more thinking than the task needs, or a new MCP server, hook or plugin is adding work. Run /context and /usage first, then claude --safe-mode to rule out your setup.

How do I make Claude Code faster?

Run /clear between unrelated tasks, lower the effort level with /effort low for simple work, switch to a lighter model with /model, and disable MCP servers you are not using. If you are on Opus and can pay for it, /fast makes responses up to 2.5x faster at a higher per-token price.

Why is the first message after a break so slow?

Claude Code caches your conversation prefix so later turns are cheap and quick. That cache lasts an hour on a subscription and five minutes on an API key or once you draw on usage credits. After it expires, your next message reprocesses the whole context, which is slow on a long session. Starting fresh with /clear or resuming from a summary avoids paying for that.

Does a big CLAUDE.md slow Claude Code down?

A little, on every turn, because it loads into context at session start. Anthropic's guidance is to keep CLAUDE.md under 200 lines and move workflow-specific instructions into skills, which load only when invoked. A CLAUDE.md that size is rarely the main cause of a slow session, but it is the cheapest thing to trim.

Is Claude Code slow because of Anthropic's servers?

Sometimes. If you see API Error 529 Overloaded or repeated retries, the slowness is on the service side and no local setting will fix it. If claude --safe-mode is just as slow and /context shows a small session, check Anthropic's status page before you spend time tuning your own setup.

Keep it fast by default

Most slow sessions are fixed by habits, not settings: clear between tasks, keep effort matched to the job, and keep the always-on setup small. That last part is what ClockedCode is for, a curated set of tools and a tight CLAUDE.md so a fresh session starts lean.