GPT-5.4 Mini API
The compact GPT-5.4 at $0.75 / $4.50 — fast and cheap for volume work on the OpenAI-compatible endpoint.
Searching for the GPT Mini API? GPT-5.4 Mini is the current GPT Mini release — same endpoint and key across every version, so code written for the GPT Mini API works unchanged when new versions ship.
Pricing
| Rate | Official | AI Prime Tech | You save |
|---|---|---|---|
| Input (per 1M tokens) | $0.75 | $0.1 | 87% OFF |
| Output (per 1M tokens) | $4.5 | $0.59 |
Pay-as-you-go, no subscription. Prompt caching is passed through from the upstream model.
GPT-5.4 Mini API pricing in detail
Anthropic pricing for GPT-5.4 Mini is per million input tokens and per million output tokens: $0.75 and $4.5 at list. Output tokens are the expensive side, and long answers, extended thinking and agent loops are output-heavy — which is why a model that gets the task right in one attempt is usually cheaper than a nominally cheaper one that needs three. On this gateway the same list rates are billed in credits that cost 87% less than face value, so the effective price is $0.1 per million input tokens and $0.59 per million output tokens.
Prompt caching. A long, stable system prompt or document can be cached; cache writes cost 25% more than a normal input token, and cache hits (cache reads) about a tenth of one, so a 20k-token system prompt reused across a session costs a fraction of resending it. The gateway forwards the cache_control blocks unchanged, and cache read tokens show separately in your usage log.
Batch processing. Anthropic’s Batch API prices asynchronous jobs at half the real-time rate. The gateway serves the real-time Messages API only — batch endpoints are not proxied — so if half-price overnight batches are the core of your workload, run those directly on Anthropic and keep the interactive traffic here, where the credit discount is larger than the batch discount anyway.
Worked example. A Claude Code session that sends 2 million input tokens and receives 100k output tokens costs $1.95 at Anthropic list price and $0.26 here; with a cached system prompt the input side drops further.
Best for
- Bulk generation
- Chat backends
- Extraction
- Cost-sensitive tools
Specifications
| Release | March 2026 |
| Context window | 272K tokens |
| Max output | 128K tokens |
| Vision | Yes |
| Tool use | Yes |
| Model ID | gpt-5.4-mini |
Quick start
OpenAI SDK, OpenAI-compatible endpoint. Create your API key in the Codex group, then point the SDK at the gateway and use model ID gpt-5.4-mini:
export OPENAI_BASE_URL=https://aiprimetech.io/v1
export OPENAI_API_KEY=sk-your-codex-group-keyWorks with Codex CLI, Cursor, Cline, Continue and any OpenAI-SDK codebase (chat completions and Responses API).
Works with Claude Code, Cursor, Cline, Aider, and the official Anthropic / OpenAI SDKs. See the 60-second setup guide.
Instant API key, pay-as-you-go, crypto accepted. 87% below official pricing.
Get Your API Key →Frequently asked questions
How expensive is GPT-5.4 Mini API?
GPT-5.4 Mini lists at $0.75 per million input tokens and $4.5 per million output tokens on Anthropic pricing. Through this gateway the same calls cost $0.1 / $0.59, because credits are bought 87% below face value. Cache hits are cheaper still.
What is the model id for GPT-5.4 Mini?
gpt-5.4-mini — pass it as model in the Messages API or select it with /model in Claude Code. It is listed by GET /v1/models on this gateway.
Does the GPT-5.4 Mini API support prompt caching and batch processing here?
Prompt caching yes — cache_control is forwarded and cache reads are billed at the reduced rate. The Batch API is not proxied; interactive, real-time requests only.
Is this the official Anthropic API?
No. AI Prime Tech is an independent gateway that serves GPT-5.4 Mini through the same Messages API shape and model id. Nothing about the model changes; the billing does.
All models: Claude Fable 5.1 · Claude Opus 5.5 · Claude Sonnet 5.5 · Claude Opus 5 · Claude Sonnet 5 · Claude Fable 5 · Claude Opus 4.8 · Claude Opus 4.7 · Claude Opus 4.6 · Claude Sonnet 4.6 · Claude Opus 4.5 · Claude Sonnet 4.5 · Claude Haiku 4.5 · Claude 1M Context · GPT-6 Astra · GPT-6 Sol · GPT-6 Luna · GPT-5.6 · GPT-5.6 Sol · GPT-5.6 Terra · GPT-5.6 Luna · GPT-5.5 · GPT-5.4 · GPT-5.3 Codex Spark · GPT-5.2 · GPT-5.2 Pro · Gemini 3 Pro · Gemini 3 Flash · MiniMax M3