Early build · scores computed by the deterministic engine from real, sourced developments (30 days) · weights v1.0-prior
LatestModelsCompare › Longest context
Vikshy Compare

AI models with the longest context windows

The AI models that can hold the most text at once, largest context window first. The largest listed is Llama 4 Scout (Meta) at up to 10M tokens. A larger window lets a model reason over more of your documents, code, or conversation in a single call.

ModelProviderContextInput /MOutput /MAccess
Llama 4 ScoutEfficientVery long context tasks and efficient on device or single GPU deployment.
Meta
10M
Open
Open
Open weights + third-party API
MiniMax-Text-01OpenExtreme long-context research and self-hosting on an open linear-attention model.
MMiniMax
4M
Open
Open
API (model id: MiniMax-Text-01) via platform.minimax.io, plus open weights on Hugging Face and GitHub.
Gemini 3.1 ProFrontierComplex analysis over very long documents. Rates rise above 200K tokens.
Google
2M
$2.00
$12.00
API + Gemini app
Grok 4.20ReasoningMulti-agent and long-horizon jobs.
xAI
2M
$1.25
$2.50
API
Grok 4.1 FastEfficientHigh-volume, cost-sensitive tasks.
xAI
2M
$0.20
$0.50
API
GPT-5.6 SolFrontierTeams tackling the hardest agentic and coding work.
OpenAI
1M
$5.00
$30.00
API + ChatGPT
GPT-5.6 TerraBalancedEveryday production apps that need strong quality on a budget.
OpenAI
1M
$2.50
$15.00
API + ChatGPT
GPT-5.6 LunaEfficientChat, classification, and other high-throughput jobs.
OpenAI
1M
$1.00
$6.00
API + ChatGPT
GPT-5.5FlagshipBroad production use across chat and tools.
OpenAI
1M
$5.00
$30.00
API + ChatGPT
GPT-5.5 ProReasoningResearch-grade analysis where accuracy beats cost.
OpenAI
1M
$30.00
$180.00
API + ChatGPT
GPT-5.4 miniEfficientLightweight tasks and cost-sensitive apps.
OpenAI
1M
$0.75
$4.50
API + ChatGPT
GPT-5.4 nanoEfficientClassification, extraction, and routing at scale.
OpenAI
1M
$0.20
$1.25
API
Gemini 3.6 FlashBalancedFast, affordable agentic and coding work.
Google
1M
$1.50
$7.50
API + Gemini app
Gemini 3.5 FlashBalancedGeneral production workloads.
Google
1M
$1.50
$9.00
API + Gemini app
Gemini 3.5 Flash-LiteEfficientLarge-scale, latency-sensitive tasks.
Google
1M
$0.30
$2.50
API
Gemini 3.5 Flash CyberEnterpriseSecurity teams. No public API pricing or self-serve access.
Google
1M
Enterprise
Claude Fable 5FrontierAmbitious, long-horizon agent work.
Anthropic
1M
$10.00
$50.00
API + Claude apps
Claude Mythos 5EnterpriseApproved security teams in Project Glasswing. Invitation only.
Anthropic
1M
$10.00
$50.00
Limited availability
Claude Opus 4.8FlagshipComplex agentic coding and enterprise work.
Anthropic
1M
$5.00
$25.00
API + Claude apps
Claude Sonnet 5BalancedMost production workloads. Intro pricing of $2 in, $10 out runs through Aug 31, 2026.
Anthropic
1M
$3.00
$15.00
API + Claude apps

Common questions

Which AI model has the largest context window?

Of the models Vikshy tracks, Llama 4 Scout (Meta) has the largest listed context window, up to 10M tokens. Context is how much text (roughly, words plus code plus formatting) a model can consider at once.

Does a bigger context window mean a better model?

Not on its own. A large window helps with long documents and codebases, but capability, latency, and price still matter. Compare those in the table above.

← Back to the full model catalogue