Early build · scores computed by the deterministic engine from real, sourced developments (30 days) · weights v1.0-prior
LatestModelsCompare › Cheapest
Vikshy Compare

Cheapest AI models by API price

The AI models with the lowest API input price across every lab Vikshy tracks, cheapest first. 4 offer a rate-limited free tier, and the lowest per-token paid price is Command R7B (Cohere) at $0.0375 per million input tokens. Ranked purely by listed input price; open-weight models (which you self-host) are covered in open-source models.

ModelProviderContextInput /MOutput /MAccess
GLM-4.7-FlashEfficientHigh-volume, latency-sensitive tasks and prototyping at no token cost.
ZZhipu AI (GLM)
128K
Free
Free
API (model id: glm-4.7-flash) via Z.ai, free tier.
GLM-4.6V-FlashEfficientOn-device or high-volume multimodal tasks at no token cost.
ZZhipu AI (GLM)
128K
Free
Free
API (model id: glm-4.6v-flash) via Z.ai, free tier, plus open weights on Hugging Face.
GLM-4.5-FlashEfficientPrototyping and high-volume light tasks at no token cost.
ZZhipu AI (GLM)
131K
Free
Free
API (model id: glm-4.5-flash) via Z.ai, free tier.
Hunyuan-LiteEfficientPrototyping and high-volume light workloads at no token cost.
TTencent (Hunyuan)
256K
Free
Free
API (model id: hunyuan-lite) via Tencent Cloud, free tier.
Command R7BEfficientHigh volume lightweight generation.
CCohere
128K
$0.0375
$0.15
API + open weights
Qwen-FlashEfficientSimple tasks that need fast cheap responses.
Alibaba (Qwen)
1M
$0.05
$0.40
API
Devstral SmallOpenSelf hosted coding agents on a budget.
Mistral AI
256K
$0.10
$0.30
API + open weights
Hunyuan-TurboSFlagshipHigh-throughput general chat and assistant tasks that need speed and price efficiency.
TTencent (Hunyuan)
256K
$0.11
$0.28
API (model id: hunyuan-turbos) via Tencent Cloud.
Embed 4EnterpriseEnterprise semantic search and multimodal retrieval.
CCohere
128K
$0.12
N/A
API
DeepSeek-V4-FlashBalancedEveryday chat, coding and high volume production use.
DeepSeek
1M
$0.14
$0.28
API + open weights
Hunyuan Hy3FrontierCost-efficient frontier reasoning and agent work, positioned against other top open models.
TTencent (Hunyuan)
256K
$0.14
$0.58
API (model id: hy3) via Tencent Cloud TokenHub, plus open weights on Hugging Face and GitHub. Also in Yuanbao, CodeBuddy, and ima.
Pixtral 12BOpenImage plus text apps that want open weights.
Mistral AI
128K
$0.15
$0.15
API + open weights
Command RLegacyCost sensitive RAG and chat.
CCohere
128K
$0.15
$0.60
API + open weights
MiniMax-M2.5BalancedHigh-volume coding and agent work where value per token is the priority.
MMiniMax
204K
$0.15
$0.90
API (model id: MiniMax-M2.5) via platform.minimax.io.
GPT-5.4 nanoEfficientClassification, extraction, and routing at scale.
OpenAI
1M
$0.20
$1.25
API
Grok 4.1 FastEfficientHigh-volume, cost-sensitive tasks.
xAI
2M
$0.20
$0.50
API
moonshot-v1-8kLegacySimple, cheap legacy calls that do not need long context.
KMoonshot AI (Kimi)
8K
$0.20
$2.00
API only (model id: moonshot-v1-8k). Hosted.
Gemini 3.5 Flash-LiteEfficientLarge-scale, latency-sensitive tasks.
Google
1M
$0.30
$2.50
API
CodestralEfficientIDE autocomplete and high frequency code generation.
Mistral AI
256K
$0.30
$0.90
API
GLM-4.6VBalancedMultimodal reasoning and document or UI understanding with tool use.
ZZhipu AI (GLM)
128K
$0.30
$0.90
API (model id: glm-4.6v) via Z.ai, plus open weights on Hugging Face.

Common questions

What is the cheapest AI model?

The lowest-cost options are free tiers (rate-limited): GLM-4.7-Flash (Zhipu AI (GLM)), GLM-4.6V-Flash (Zhipu AI (GLM)), GLM-4.5-Flash (Zhipu AI (GLM)). Among models with a per-token price, Command R7B (Cohere) is the lowest at $0.0375 per million input tokens. Cheaper is not better for every task; check the context window and what the model is built for.

Are open-weight models cheaper?

Open-weight models have no per-token API price because you run them yourself, so hosting cost replaces API cost. See the open-source models list.

← Back to the full model catalogue