Zed-Hosted Models

Zed's plans offer hosted versions of major LLMs with higher rate limits than direct API access. Model availability is updated regularly. To use your own API keys instead, see LLM Providers. For general setup, see AI Quick Start.

Note: Claude Opus models, GPT-5.5 pro, and GPT-5.4 pro are only available on Zed Pro and Zed Business.

Model	Provider	Token Type	Provider Price per 1M tokens	Zed Price per 1M tokens
Claude Opus 4.5	Anthropic	Input	$5.00	$5.50
	Anthropic	Output	$25.00	$27.50
	Anthropic	Input - Cache Write	$6.25	$6.875
	Anthropic	Input - Cache Read	$0.50	$0.55
Claude Opus 4.6	Anthropic	Input	$5.00	$5.50
	Anthropic	Output	$25.00	$27.50
	Anthropic	Input - Cache Write	$6.25	$6.875
	Anthropic	Input - Cache Read	$0.50	$0.55
Claude Opus 4.7	Anthropic	Input	$5.00	$5.50
	Anthropic	Output	$25.00	$27.50
	Anthropic	Input - Cache Write	$6.25	$6.875
	Anthropic	Input - Cache Read	$0.50	$0.55
Claude Opus 4.8	Anthropic	Input	$5.00	$5.50
	Anthropic	Output	$25.00	$27.50
	Anthropic	Input - Cache Write	$6.25	$6.875
	Anthropic	Input - Cache Read	$0.50	$0.55
Claude Sonnet 4.5	Anthropic	Input	$3.00	$3.30
	Anthropic	Output	$15.00	$16.50
	Anthropic	Input - Cache Write	$3.75	$4.125
	Anthropic	Input - Cache Read	$0.30	$0.33
Claude Sonnet 4.6	Anthropic	Input	$3.00	$3.30
	Anthropic	Output	$15.00	$16.50
	Anthropic	Input - Cache Write	$3.75	$4.125
	Anthropic	Input - Cache Read	$0.30	$0.33
Claude Haiku 4.5	Anthropic	Input	$1.00	$1.10
	Anthropic	Output	$5.00	$5.50
	Anthropic	Input - Cache Write	$1.25	$1.375
	Anthropic	Input - Cache Read	$0.10	$0.11
GPT-5.5 pro	OpenAI	Input	$30.00	$33.00
	OpenAI	Output	$180.00	$198.00
GPT-5.5	OpenAI	Input	$5.00	$5.50
	OpenAI	Output	$30.00	$33.00
	OpenAI	Cached Input	$0.50	$0.55
GPT-5.4 pro	OpenAI	Input	$30.00	$33.00
	OpenAI	Output	$180.00	$198.00
GPT-5.4	OpenAI	Input	$2.50	$2.75
	OpenAI	Output	$15.00	$16.50
	OpenAI	Cached Input	$0.025	$0.0275
GPT-5.3-Codex	OpenAI	Input	$1.75	$1.925
	OpenAI	Output	$14.00	$15.40
	OpenAI	Cached Input	$0.175	$0.1925
GPT-5.2	OpenAI	Input	$1.75	$1.925
	OpenAI	Output	$14.00	$15.40
	OpenAI	Cached Input	$0.175	$0.1925
GPT-5.2-Codex	OpenAI	Input	$1.75	$1.925
	OpenAI	Output	$14.00	$15.40
	OpenAI	Cached Input	$0.175	$0.1925
GPT-5 mini	OpenAI	Input	$0.25	$0.275
	OpenAI	Output	$2.00	$2.20
	OpenAI	Cached Input	$0.025	$0.0275
GPT-5 nano	OpenAI	Input	$0.05	$0.055
	OpenAI	Output	$0.40	$0.44
	OpenAI	Cached Input	$0.005	$0.0055
Gemini 3.1 Pro	Google	Input	$2.00	$2.20
	Google	Output	$12.00	$13.20
Gemini 3.5 Flash	Google	Input	$1.50	$1.65
	Google	Output	$9.00	$9.90
Gemini 3 Flash	Google	Input	$0.50	$0.55
	Google	Output	$3.00	$3.30

Recent Model Retirements

As of February 19, 2026, Zed Pro serves newer model versions in place of the retired models below:

Claude Opus 4.1 → Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, or Claude Opus 4.8
Claude Sonnet 4 → Claude Sonnet 4.5 or Claude Sonnet 4.6
Claude Sonnet 3.7 (retired Feb 19) → Claude Sonnet 4.5 or Claude Sonnet 4.6
GPT-5.1 and GPT-5 → GPT-5.2 or GPT-5.2-Codex
Gemini 2.5 Pro → Gemini 3.1 Pro
Gemini 3 Pro → Gemini 3.1 Pro
Gemini 2.5 Flash → Gemini 3 Flash or Gemini 3.5 Flash

Usage

Any usage of a Zed-hosted model will be billed at the Zed Price (rightmost column above). See Plans & Pricing for details on Zed's plans and limits for use of hosted models.

Because Zed-hosted Gemini models do not use Google context caching, Gemini usage is billed only as input and output tokens; there is no separate cached-input price for these models. This preserves zero-data-retention behavior for hosted Gemini requests. For background, see Google's Vertex AI documentation on context caching and zero data retention.

LLMs can enter unproductive loops that require user intervention. Monitor longer-running tasks and interrupt if needed.

Context Windows

A context window is the maximum span of text and code an LLM can consider at once, including both the input prompt and output generated by the model.

Model	Provider	Zed-Hosted Context Window
Claude Opus 4.5	Anthropic	200k
Claude Opus 4.6	Anthropic	1M
Claude Opus 4.7	Anthropic	1M
Claude Opus 4.8	Anthropic	1M
Claude Sonnet 4.5	Anthropic	200k
Claude Sonnet 4.6	Anthropic	1M
Claude Haiku 4.5	Anthropic	200k
GPT-5.5 pro	OpenAI	272k input / 400k total
GPT-5.5	OpenAI	272k input / 400k total
GPT-5.4 pro	OpenAI	272k input / 400k total
GPT-5.4	OpenAI	272k input / 400k total
GPT-5.3-Codex	OpenAI	272k input / 400k total
GPT-5.2	OpenAI	272k input / 400k total
GPT-5.2-Codex	OpenAI	272k input / 400k total
GPT-5 mini	OpenAI	272k input / 400k total
GPT-5 nano	OpenAI	272k input / 400k total
Gemini 3.1 Pro	Google	200k
Gemini 3.5 Flash	Google	1M
Gemini 3 Flash	Google	1M

Zed currently limits hosted Gemini 3.1 Pro requests to 200k tokens because pricing changes above that context size.

Each Agent thread in Zed maintains its own context window. The more prompts, attached files, and responses included in a session, the larger the context window grows.

Start a new thread for each distinct task to keep context focused.

Tool Calls

Models can use tools to interface with your code, search the web, and perform other useful functions.