· 8 min · Dev Guides

CLAUDE_CODE_MAX_OUTPUT_TOKENS Explained: Output Limits and Plugin Folders in Claude Code 2.1.265

CLAUDE_CODE_MAX_OUTPUT_TOKENS Explained: Output Limits and Plugin Folders in Claude Code 2.1.265

claude_code_max_output_tokens is the search phrase for CLAUDE_CODE_MAX_OUTPUT_TOKENS, the Claude Code environment variable that sets the maximum output tokens for most requests. This guide explains the limits verified by September 10, 2026, how configuration examples should be interpreted, and what Claude Code 2.1.265 changes about plugin folders and saved tool results.

What claude_code_max_output_tokens controls

The environment-variable reference describes CLAUDE_CODE_MAX_OUTPUT_TOKENS as setting “the maximum number of output tokens for most requests.” That wording matters: it describes a request output budget, with defaults and caps that vary by model.

The variable is therefore relevant when investigating a Claude Code output limit. It should not be treated as a universal limit for every kind of data Claude Code handles. Version 2.1.265 also introduces a disk limit for saved tool results, but that limit measures stored tool data in gigabytes, not generated output in tokens.

There is a historical qualification to the configuration documentation. The environment-variable reference is continuously updated, and its exact September 10 revision was not recovered. Its description explains the variable’s purpose, but the page does not establish every default that applied to Claude Code 2.1.265.

For a dated article, the strongest numerical evidence comes from version-tagged release notes. Those notes establish specific earlier changes without requiring a complete default table to be inferred from the live documentation.

Three limits to keep separate

Limit What it concerns Verified information available by the cutoff
Request output budget Maximum output tokens for most requests Model-specific changes documented in v2.1.77
Model context window The model’s documented context capacity A 1-million-token window for Sonnet 5, Opus 5, and Fable 5.1
Saved tool-result cap Tool results saved to disk A 1 GB cap introduced in v2.1.265

When diagnosing a limit, first identify which of these applies. A token-budget setting and a disk-storage cap require different interpretations, even when both appear during the same development session.

Verified defaults and maximum output limits

The v2.1.77 release notes, dated March 17, 2026, document two relevant changes: the Claude Opus 4.6 default increased to 64k output tokens, and the upper bound for Opus 4.6 and Sonnet 4.6 increased to 128k tokens. These changes are also preserved in the v2.1.265-tagged changelog.

That evidence rules out presenting “32,000 default, 64,000 maximum” as a universal Claude Code rule. It also does not justify filling in missing defaults for other models.

Model Claude Code default established by the dated sources Documented upper bound
Claude Opus 4.6 Increased to 64k in v2.1.77 128k in v2.1.77
Claude Sonnet 4.6 Not established by the dated sources reviewed 128k in v2.1.77

The careful distinction is between a documented change and a complete configuration reference. The March release establishes the changes above; the available dated sources do not establish a full model-by-model default table for v2.1.265.

Newer model specifications do not establish CLI defaults

The model release notes document a 128k maximum output and a 1-million-token context window for Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1. Their exact API IDs are claude-sonnet-5, claude-opus-5, and claude-fable-5-1, respectively.

Those announcements were available before the publication cutoff: Sonnet 5 on June 30, Opus 5 on July 24, and Fable 5.1 on September 1, 2026. However, their model specifications are not evidence of Claude Code’s default request budget.

For developers choosing a model, the Claude Code models guide provides related reading. Keep the model’s published maximum separate from the default output budget used by the CLI.

Setting the variable and understanding compaction

The live documentation supports setting the variable before launching claude, or placing it under the env key in settings.json. Because the exact documentation revision available on September 10 was not recovered, the examples below are documented configuration patterns with that historical qualification.

Shell configuration

The supplied shell example sets the output budget to 64,000 tokens:

export CLAUDE_CODE_MAX_OUTPUT_TOKENS=64000

The documented pattern is to set the variable before launching Claude Code. The number in this example is a chosen configuration value; it is not a claim that 64,000 is the default or maximum for every model.

When recording configuration for a project or troubleshooting session, identify the selected model alongside the requested budget. Without the model, a value alone cannot explain whether it matches a documented default or falls within a documented upper bound.

Configuration in settings.json

The equivalent supplied JSON example places the variable under env:

{"env":{"CLAUDE_CODE_MAX_OUTPUT_TOKENS":"64000"}}

The value is a string in this example. As with the shell setting, the snippet expresses a requested budget rather than proving a model-specific default.

For related command-line usage, see the Claude Code CLI reference. The historical evidence here supports the variable’s role and the supplied configuration examples, while leaving the complete v2.1.265 defaults unresolved.

What is verified about compaction?

The v2.1.265-tagged changelog records an earlier fix in v2.1.69 for CLAUDE_CODE_MAX_OUTPUT_TOKENS being ignored during conversation compaction. That is a verified connection between this setting and compaction behavior.

The live environment-variable reference also explains that increasing the setting reduces the effective context available before auto-compaction. The publication date of that explanation is uncertain, so it cannot be presented as a verified September 10 description.

The setting should therefore not be described as enlarging the model’s published context window. The dated evidence establishes an output-budget control and a compaction-related fix; it does not establish a precise compaction threshold for every model in v2.1.265.

How Claude Code 2.1.265 loads plugin folders

Claude Code v2.1.265 was released on September 8, 2026, with release commit 8e02f6d. Its folder-loading change allows --plugin-dir to point at a folder containing multiple plugins.

The supplied invocation is:

claude --plugin-dir ./plugins

The release describes the discovery rule directly: “each child folder with a manifest loads”. It also states that additions and removals are detected while Claude Code runs.

This makes the parent folder the useful unit when managing several plugins together. The documented contract concerns child folders that contain manifests; the release notes do not specify recursive discovery. Do not assume that an arbitrarily nested directory tree will be searched.

Manifest layout and the historical caveat

The live plugin reference places the manifest at .claude-plugin/plugin.json and component directories such as skills/, commands/, and hooks/ at the plugin root. That reference was not recovered as a dated v2.1.265 snapshot.

For the September 10 article, distinguish that live layout guidance from the release’s explicit guarantee: child folders with manifests load. The release establishes multi-plugin folder discovery without establishing every detail of the contemporaneous layout documentation.

Plugin fixes included in the release

Version 2.1.265 also fixes three plugin issues:

  • A plugin path containing backslashes could bypass symlink containment on macOS and Linux.
  • Directories beginning with two dots were incorrectly rejected.
  • Unreadable default component folders were silently skipped; they now appear in /plugin with an error code.

These fixes concern plugin paths and component loading. If a plugin folder fails to load, they provide relevant release context, but the output-token setting does not explain those directory-handling issues.

The 1 GB saved tool-result cap is a separate limit

Version 2.1.265 adds a 1 GB cap on tool results saved to disk. The conversation preview reports truncation.

This limit concerns stored tool data. CLAUDE_CODE_MAX_OUTPUT_TOKENS concerns maximum output tokens for most requests. The release does not describe the environment variable as a control for the disk cap, so increasing its value should not be presented as a remedy for saved tool-result truncation.

The release notes also do not specify whether the 1 GB cap applies per result, per session, or across another grouping. Preserve that uncertainty when documenting behavior: “1 GB cap on tool results saved to disk” is the verified description.

A useful troubleshooting distinction is the evidence shown to the developer. A conversation preview reporting truncation relates to the new saved-result behavior. A question about a model’s generated output budget belongs with the selected model and token setting. A plugin error code in /plugin belongs with plugin loading.

These are separate controls and diagnostics introduced or discussed in the same release context. Treating them separately prevents a token configuration change from being used to explain unrelated storage or folder behavior.

The gateway follow-up available by September 10

Version 2.1.266 was also released on September 8, so it falls within this article’s cutoff. The v2.1.266 release notes fix a v2.1.265 gateway-authentication regression involving CLAUDE_CODE_USE_GATEWAY.

The relevant error is “Not signed in to the Cloud gateway”. The release states that no configuration change is needed.

For a developer encountering that exact regression on v2.1.265, the follow-up release is directly relevant. Raising the output-token budget or reorganizing plugin folders is not the documented fix.

The same workflows can run through an API gateway such as AI Prime Tech; see using Claude Code with an API key for related setup guidance. AI Prime Tech operates independently and is not affiliated with Anthropic or OpenAI.

Key takeaways

  • CLAUDE_CODE_MAX_OUTPUT_TOKENS sets the maximum output tokens for most requests; defaults and caps vary by model.
  • v2.1.77 increased Opus 4.6’s default to 64k and the Opus 4.6/Sonnet 4.6 upper bound to 128k. A complete v2.1.265 default table remains unverified.
  • Model maximum output, model context capacity, and saved tool-result storage are separate limits.
  • In v2.1.265, --plugin-dir can load multiple plugins from child folders with manifests and detect additions and removals while running.
  • v2.1.265 introduces the 1 GB saved tool-result cap; v2.1.266 fixes the gateway-authentication regression with no configuration change needed.

FAQ

What does CLAUDE_CODE_MAX_OUTPUT_TOKENS control?

CLAUDE_CODE_MAX_OUTPUT_TOKENS sets the maximum number of output tokens for most requests. Its defaults and caps vary by model, and it is separate from v2.1.265’s 1 GB cap on tool results saved to disk.

What are the default and maximum output limits?

v2.1.77 increased Claude Opus 4.6’s default to 64k output tokens and the Opus 4.6/Sonnet 4.6 upper bound to 128k. The dated sources reviewed do not establish Sonnet 4.6’s default or a complete model-by-model default table for v2.1.265.

Does raising the setting increase the context window?

The documented purpose is to set an output-token budget, not to increase the model’s published context window. v2.1.69 fixed the variable being ignored during compaction, but the live documentation’s explanation of its effect on auto-compaction has an uncertain publication date.

How does --plugin-dir discover multiple plugins?

In Claude Code 2.1.265, --plugin-dir can point to a parent folder, and each child folder with a manifest loads. Additions and removals are detected while Claude Code runs; recursive discovery is not specified in the release notes.

A

AI Prime Tech publishes this blog. Articles are drafted with AI assistance; check version-specific details such as model IDs, prices and limits against the official documentation before relying on them.

Get cheaper Claude API access

One API key for Claude Opus 5.5, Sonnet 5, Haiku 4.5 and Fable 5.1, plus GPT-6 models. Pay as you go, no subscription.

Get Your API Key →
AI Prime Tech is an independent third-party API gateway. Claude™ and Anthropic® are trademarks of Anthropic, PBC. No affiliation or endorsement is implied.