Direct-API and provider pricing for current frontier & value LLMs — USD per 1M tokens.
| Model | Tier | Context | Max out | Best price ($/M in / out) | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Flashcore
DeepSeek |
cheap | 1M | 384K | $0.07 / $0.11 via InferX · 4 providers |
Efficiency-optimized MoE (284B/13B active). 1M ctx, thinking mode. |
| Nube-Choice
Nube |
cheap | 1M | 384K | $0.154 / $0.462 via Nube · 1 provider |
Nube's auto-routing model. Currently routes to DeepSeek-V4-Flash (price tracks the routed model; subject to change as routing shifts). |
| DeepSeek V4 Procore
DeepSeek |
frontier | 1M | 384K | $1.32 / $3.96 via Direct API · 1 provider |
DeepSeek's flagship. Same 1M ctx / 384K out, higher ceiling. |
| Qwen3.8-Maxcore
Alibaba Qwen |
frontier | 1M | 131K | $2.00 / $6.00 via Direct API · 1 provider |
2.4T MoE (95B active), multimodal. Flagship of Qwen3.8 series. |
| Qwen3.7-Pluscore
Alibaba Qwen |
mid | 1M | 131K | $0.32 / $1.28 via Direct API · 1 provider |
Cost-effective workhorse in the Qwen3.7 series. |
| Qwen3.7-Flash
Alibaba Qwen |
cheap | 1M | 131K | $0.03 / $0.13 via Direct API · 1 provider |
Qwen's budget text tier, 1M ctx. Cheapest model on the board. |
| Qwen3.8-27B
Alibaba Cloud (Qwen) |
mid | 0.262M | 262K | $0.175 / $1.05 via Nube · 2 providers |
Open-weight (Apache-2.0) dense VLM, 262K ctx. Served via Cloudflare/OpenRouter/Vercel. |
| GPT-5.6 Luna
OpenAI |
cheap | 1.05M | 128K | $0.2 / $1.20 via Direct API · 1 provider |
OpenAI's cheapest GPT-5.6 tier. Fast, low-cost general model. |
| GPT-5.6 Terra
OpenAI |
mid | 1.05M | 128K | $2.00 / $12.00 via Direct API · 1 provider |
Mid-tier GPT-5.6 (Terra). Balances cost and capability. |
| Claude Opus 5
Anthropic |
frontier | 1M | 64K | $5.00 / $25.00 via Direct API · 1 provider |
Anthropic's frontier reasoning model. |
| Claude Sonnet 5
Anthropic |
mid | 1M | 64K | $2.00 / $10.00 via Direct API · 1 provider |
Balanced mid-tier Claude for agents and tool use. |
| GLM-5.3
Zhipu AI |
frontier | 1M | — | $0.49 / $1.54 via Nube · 2 providers |
Zhipu flagship coding/cyber model, 1M ctx. |
| GLM-5.3-Flash
Zhipu AI |
cheap | 1M | — | $0.06 / $0.2 via Nube · 2 providers |
Natively-multimodal (first in GLM-5 series), 320B/18B active, 1M ctx. |
| MiniMax M3
MiniMax |
cheap | 1M | — | $0.3 / $1.20 via Direct API · 1 provider |
Multimodal foundation model; very cheap per token. |
| MiMo-V2.5
Xiaomi |
cheap | 1M | 131K | $0.14 / $0.28 via Direct API · 1 provider |
Xiaomi's omnimodal LLM (image/video/audio/text), 1M ctx. |
| Qwen3.7-Max
Alibaba Qwen |
frontier | 1M | 131K | $1.25 / $3.75 via Direct API · 1 provider |
Previous Qwen flagship; promotional pricing active. |
| GPT-5.6 Sol
OpenAI |
frontier | 1.05M | 128K | $5.00 / $30.00 via Direct API · 1 provider |
OpenAI's top GPT-5.6 tier. |
| Kimi K3
Moonshot AI |
frontier | 1M | — | $1.65 / $8.25 via Nube · 2 providers |
Newest Kimi flagship, 1M ctx. |
| Model | Vendor | Est. monthly cost | Cost / 1M out | Tier |
|---|
Beyond the four requested models, these are the closest current comparables (already in the table above):
Direct = the vendor's own API list price. Provider columns (Nube / Fireworks / InferX) = each provider's posted rate for that model. Prices are USD per 1M tokens; DeepSeek shown at peak with cache-hit input pricing — off-peak is 50% lower. Cache-hit discounts and batch pricing vary by vendor. Promotional prices are marked. Figures were hand-verified against official docs and provider pricing pages on the generation date.