prompt is too long: 209117 tokens > 200000 maximum
This is a rejection, not a failure. The request was measured and refused before the model saw any of it, so nothing was processed, nothing was truncated, and nothing was billed as work. It is not a rate limit, not a quota, and not an outage — those all involve a request that was allowed through. The first number is what you sent; the second is the ceiling you sent it against.
Data as of 2026-09. Vendor limits and defaults change; check the official docs for current values before acting on any number below.
Read the second number first
The count on the left changes every time and tells you almost nothing. The number on the right identifies which context window you are actually running with, and that is frequently a surprise.
If you believe you are on a model with a 1M-token window and the ceiling reads
200000, you are not running with that window. Documented ways to end up at
200K: Claude Haiku 4.5’s context window is 200K; Opus 4.8 and Opus 5 run with a
200K window on Amazon Bedrock, Google Cloud’s Agent Platform and Microsoft
Foundry; and CLAUDE_CODE_DISABLE_1M_CONTEXT=1 puts a native-1M model on the
200K boundary. Before you spend an hour trimming context, confirm the ceiling is
the one you expect.
Also note where you are seeing this exact wording. An interactive Claude Code
session renders the same condition as Context limit reached · /compact or /clear to continue. The raw prompt is too long form is what appears in -p
output and in the transcript — so if you are reading this string, you are most
likely in a headless run, a script, a subagent, or an SDK app, where there is no
interactive /compact to bail you out. Amazon Bedrock words the same condition
Input is too long for requested model., and a Claude apps gateway reports it as
capability_rejected: prompt_too_long.
The floor that compaction cannot go below
Here is the part that changes what you do about it. Everything in the request
shares one budget: the system prompt, CLAUDE.md and memory, MCP tool
definitions, skills, attachments, files that were read — and only then your
conversation.
Those non-conversation pieces are re-injected after every compaction. The
system prompt, CLAUDE.md, memory and MCP tool definitions come back; the body
of each skill you invoked comes back capped at 5,000 tokens per skill; Claude
Code re-reads up to five of the files touched in the session, with any file over
5,000 tokens returning as a path reference instead of its contents. Compaction
shrinks the conversation down onto that baseline. It cannot shrink the baseline
itself.
So a session can be structurally uncompactable. Claude Code says this outright in
the interactive case: a single-exchange conversation cannot be compacted, because
there are no earlier turns to summarize, and the message names whether the
conversation’s own content or the system prompt, tool definitions and attachments
make up most of the request. A brand-new subagent or a fresh -p invocation that
hits this ceiling proves the same point by construction — there is no history
there to blame.
The practical version: when several MCP servers are loaded, their tool
definitions are paid for on every request, before you type anything. That is
the most common reason a -p run or a batch of subagents fails while the same
work succeeds interactively on a different machine.
Is retrying useful?
No. Nothing changes between attempts, so nothing changes about the outcome.
The request is rejected by arithmetic. The same messages tokenize to the same count and hit the same ceiling on the second attempt, the tenth, and the hundredth. A retry wrapper around this call converts one clean rejection into a loop of identical rejections, and in a scripted or CI context that loop is how people discover the error after it has run overnight.
The one retry that means anything is the one you run after /context shows a
smaller total. If the total did not move, do not send the request again.
Find the tokens before you delete anything
/context shows everything occupying the context window, broken down by
category. Open it before you take any action — the fix depends entirely on which
category is largest, and guessing wrong costs you the conversation you were
trying to keep.
It also distinguishes two conditions that look identical from the outside:
Context exceeds the 200k-token limit by 94k tokens — run /compact or /clear to continue.
Context is 94k tokens past the 200k-token compaction window — run /compact to reduce usage.
The second form means the limit you crossed is a compaction window set below
the model’s real context window, not the model’s hard limit. Requests past it can
still succeed. If that is the line you are looking at, your session is not
actually broken, and raising or lowering the window with /autocompact <value>,
the --autocompact flag, or CLAUDE_CODE_AUTO_COMPACT_WINDOW is a live option.
The accepted range is 100K to 1M tokens, capped at the model’s own window.
What to do, keyed to what /context shows
- Conversation dominates —
/compact. If/compactitself fails withConversation too long, you are in the compaction deadlock and have to remove turns before compaction can work. - MCP tool definitions dominate — remove servers this project does not use.
Scope matters: a
user-scoped server is loaded in every project you open, including the ones that never call it, whileprojectscope lives in.mcp.jsonat the project root andlocalscope in~/.claude.json. The decisive test isclaude --safe-mode, which starts with all plugins, MCP servers and hooks disabled — if the same task fits under safe mode, your customization is the floor, and no amount of compacting will help. - Files and attachments dominate — stop reading whole files. Reading by offset and limit, or searching instead of reading, keeps large files out of the window in the first place.
CLAUDE.mdand memory dominate — trimming these is a permanent saving, because they are re-injected on every compaction. Note that nested and path-scoped rule files load into message history and are summarized away with everything else, so the two behave differently.- The ceiling is not the one you expected — fix the window, not the content.
Check which platform the model is running on and whether
CLAUDE_CODE_DISABLE_1M_CONTEXTis set in your environment.
Confirming you actually fixed it
Re-run the exact request that was rejected — same prompt, same project, same flags. A shorter test prompt succeeding proves nothing, because a shorter prompt was always going to fit.
For the tool-definition case, confirm by comparing /context before and after
removing a server, not by whether the session “feels lighter”. For the headless
case, run the same -p command twice: one success can be a coincidence of which
files happened to be in context, and two in a row means the baseline genuinely
dropped below the ceiling.
Related limits that get confused with this one
- Error during compaction: Conversation too long is what you hit when you follow this page’s advice and
/compactrefuses to run — read it before you assume compaction is broken. - API Error: Claude’s response exceeded the output token maximum is the mirror image: a limit on what the model writes, not on what you send, and completely unaffected by compaction.
- Error: File content exceeds maximum allowed tokens is the guardrail that fires before a single file read can put you here.
- The context window calculator is the quickest way to see how much of the window your baseline eats before the first message.