Home › Claude Code › Errors › Prompt is too long
OverviewInstallCLI referencePricing & limitsAgent SDKSkills & pluginsGitHubWeb, desktop & IDEModelsErrorsOpen source?TutorialFor studentsQA automationvs Cursorvs Codex & OpenCodeSubagentsHooksMCPEcosystemVS CodeOpenRouterRate limits中文指南
Last updated: September 30, 2026

Claude "Prompt is too long", "input is too long for requested model" and "Context limit reached": what fills the window and how to get it back

All of these mean the same thing: the request Claude Code (or claude.ai) is about to send no longer fits the model's context window. The window holds far more than your typed messages — every file the agent read, every tool result, the system prompt, loaded skills and MCP tool definitions all count. The fix is to shrink what is sent, and the tracker shows several cases where the meter itself was wrong.

TLDR

  • Prompt is too long / input is too long for requested model — a 400 from the API: the request exceeds the window. /compact, then /clear if needed.
  • Context limit reached · /compact or /clear to continue — Claude Code's own pre-flight check; same fixes, and sometimes a false alarm.
  • automatic compaction failed / single-exchange conversation cannot be compacted — one giant turn (a huge file or tool output); /clear and re-brief with less.
  • Bulk offenders: many MCP servers, big skills, reading whole logs or minified files, long /goal texts, subagents inheriting everything.

The messages

API Error: 400 {"type":"error","error":{"type":"invalid_request_error","message":"prompt is too long: 213456 tokens > 200000 maximum"}}
API Error: 400 input is too long for requested model
Context limit reached · /compact or /clear to continue
Error during compaction: automatic compaction failed
Error: a single-exchange conversation cannot be compacted
Stop hook error: Hook evaluator API error: Prompt is too long

The first two come from the API and carry a 400 status: the request was rejected, and retrying it unchanged fails the same way. Context limit reached is Claude Code's own estimate before sending. The compaction errors mean the escape hatch itself failed, usually because the conversation is one enormous exchange with nothing older to summarise.

What actually fills the window

  • Tool results. A cat of a 4 MB log, a minified bundle, a find over node_modules — every byte of output lands in the conversation. Ask for head, grep or wc -l first, or tell the agent which files to read.
  • MCP tool definitions. Every registered server contributes its tool schemas to *every* request. #50284: the Explore subagent failed with Prompt is too long on every call once many MCP servers were registered — the subagent starts near full before it does anything. Disable servers you are not using in this project (claude mcp list, then remove or scope them).
  • Skills and CLAUDE.md. #69179 reports Context limit reached at 21% shown usage right after loading the built-in claude-api skill; another user in the thread got it on the very first prompt of a fresh process and fixed it by removing the skill invocation from CLAUDE.md. Large auto-loaded skills are a hidden fixed cost.
  • Hooks that evaluate the conversation. The /goal Stop hook sends recent conversation to an evaluator model; with a long session that evaluator request itself is what overflows (#58192, 17 comments, #62345). An Anthropic engineer noted a goal is capped at 4,000 characters, so the goal text is not the cause; the session length is. Clear the goal or compact before long autonomous runs.
  • Subagents. Each subagent has its own window, but Cowork-style spawning was reported broken with Prompt is too long at 0 tokens (#55712); update the CLI if you see that shape.

Fix it in Claude Code

/context          # what is in the window right now, by category
/compact          # summarise history, keep the task
/compact focus on the failing test and the diff   # steer what survives
/clear            # fresh window; paste a short brief of where you were
/model            # a (1M) variant if the job is genuinely large

Order matters: look at /context first. If MCP tools or a skill take a third of the window, compacting the conversation will not help for long; remove the fixed cost. If a single tool result is the problem, /clear is faster than fighting compaction. If the work legitimately needs hundreds of thousands of tokens, a (1M) model variant is the right tool — but see the next section, because selecting it has its own trap.

When the meter is wrong

Several tracker reports show the failure far below the advertised window. #51915: Context limit reached at 27% used, and /compact then errored too. #49605: the warning firing when the limit was not reached. The API Error 400 guide covers the related case of Prompt is too long landing between 166K and 178K tokens on accounts served a 200K window while the client displayed 1M. Practical rule: if /context says you have room and the API still refuses, treat about 150K as your real ceiling and compact there.

The reverse trap: picking a model whose default is the 1M window on a plan that does not include it produces a *usage-limit* message, not a context one (#62052 — "Misleading Usage limit reached error when selecting Sonnet — actually a 1M context tier gate"). Choose the non-1M variant in /model, or set CLAUDE_CODE_DISABLE_1M_CONTEXT=1. Details on the usage limit page.

"This conversation is too long to continue" on claude.ai

The chat product has the same limit with different words: *Your message will exceed the length limit for this chat. Try attaching fewer or smaller files or starting a new conversation.* and *This conversation is too long to continue. Start a new chat, or remove some tools to free up space.* The second one is about connectors: each enabled tool or MCP connector costs context in every message, which is why the message suggests removing tools (#58521). Start a new chat, attach fewer files, put stable reference material in a Project instead of re-attaching it, and switch off connectors you are not using in that conversation.

On the API and through a gateway

A 400 for length is decided by the model, so it behaves identically on Anthropic direct and through this gateway (ANTHROPIC_BASE_URL=https://aiprimetech.io). What differs is billing: the API charges per token with no session window, so a long-context job that keeps tripping subscription limits is often cheaper and calmer on a key. Check which models and window sizes the endpoint serves at /v1/models before reaching for a (1M) variant.

export ANTHROPIC_BASE_URL=https://aiprimetech.io
export ANTHROPIC_API_KEY=your_gateway_key
claude

Tired of fighting this one?

Run Claude Code through the AI Prime Tech gateway instead: the same Claude models on the gateway's own upstream accounts, two environment variables to switch, and new accounts get $5 in free API credit with code FIXIT — no card.

export ANTHROPIC_BASE_URL=https://aiprimetech.io
export ANTHROPIC_API_KEY=your_key

Create an account with code FIXIT

Frequently asked questions

What does 'input is too long for requested model' mean?

The request exceeds the context window of the model you selected. It is a 400 from the API: run /compact or /clear, remove unused MCP servers and skills, or switch to a larger-window model variant.

How do I fix 'Context limit reached' in Claude Code?

Run /context to see what fills the window, /compact to summarise, /clear if compaction fails. If it fires far below the limit, treat it as a false alarm and compact anyway; several such bugs are on the tracker.

Why does compaction fail?

Automatic compaction needs older turns to summarise. If the conversation is one giant exchange, such as a single huge tool output, there is nothing to fold; /clear and re-brief with a smaller input.

How do I fix 'This conversation is too long to continue' on claude.ai?

Start a new chat, attach fewer or smaller files, and disable connectors or tools you are not using: each enabled tool costs context in every message.

Run Claude Code on the gateway

Same models, two environment variables, credits at 7.69× face value — or a flat-rate unlimited plan. New accounts: $5 in free API credit with code FIXIT.

Get an API key See unlimited plans

AI Prime Tech is an independent API gateway and is not affiliated with, endorsed by, or sponsored by Anthropic. “Claude” and “Claude Code” are trademarks of Anthropic. Claude Code features described here follow Anthropic’s public documentation at the time of writing and change frequently; prices and model lists on this page are read from this gateway’s live settings.

More Claude Code guides

Claude Code errors and how to fix themClaude Code "API Error: 500": what it means and what actually fixes itClaude "Overloaded" error (API Error 529): what it means and what to doClaude Code API Error 400: prompt is too long, tool use concurrency, and the restClaude Code OAuth errors: timeout of 15000ms exceeded, and status code 500"Claude Code process exited with code 1": how to find the real error and fix itClaude usage limit reached: session, weekly and Opus limits, "specified API usage limits", "credit balance is too low" — and what to doClaude Code 401 and 403: invalid authentication credentials, invalid API key, "Not logged in", host_not_allowedClaude Code connection errors: ECONNRESET, connection lost mid-response, "Waiting for API response", "Unable to connect to API", SSL certificate errorsClaude Code hangs, gets stuck or stops responding: how to find out where it is stuck and get outClaude.ai not working: "Something went wrong", capacity constraints, "5-hour limit reached", "error logging you in" and the rest