The same month of AI costs $0.98 on DeepSeek vs $25 on Claude Opus, a 26× spread
We priced a typical assistant workload, 5 million input and 1 million output tokens a month, across the flagship models using live OpenRouter rates. The cheapest capable option, DeepSeek V4 Flash, runs about $0.98 a month; the priciest, Claude Opus 4.8 (batch), about $25. For most tasks a mid-tier model does the job at a fraction of the frontier price, so matching the model to the task, not defaulting to the biggest name, is where the savings are.
Method: live OpenRouter per-token rates × 5M input + 1M output tokens/month; batch and prompt-caching discounts excluded. Cite this: "usefulHQ LLM Cost Study 2026", usefulhq.com/llm-pricing/. Updates from live prices. More usefulHQ studies →
Embed or cite this stat
| Your cost/mo ▲ | Model | Provider | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|
| Free | Ling-3.0-flash (free)🧠 | Other | – | – | 262K |
| Free | Laguna S 2.1 (free)🧠 | Poolside | – | – | 262K |
| Free | Laguna XS 2.1 (free)🧠 | Poolside | – | – | 262K |
| Free | North Mini Code (free)🧠 | Cohere | – | – | 256K |
| Free | Nemotron 3.5 Content Safety (free)👁🧠 | NVIDIA | – | – | 128K |
| Free | Nemotron 3 Ultra (free)🧠 | NVIDIA | – | – | 1M |
| Free | Nemotron 3 Nano Omni (free)👁🧠 | NVIDIA | – | – | 256K |
| Free | Gemma 4 26B A4B (free)👁🧠 | – | – | 262K | |
| Free | Gemma 4 31B (free)👁🧠 | – | – | 262K | |
| Free | Lyria 3 Pro Preview👁 | – | – | 1.04858M | |
| Free | Lyria 3 Clip Preview👁 | – | – | 1.04858M | |
| Free | Nemotron 3 Super (free)🧠 | NVIDIA | – | – | 262K |
| Free | Free Models Router👁🧠 | Other | – | – | 200K |
| Free | Nemotron 3 Nano 30B A3B (free)🧠 | NVIDIA | – | – | 256K |
| Free | Nemotron Nano 12B 2 VL (free)👁🧠 | NVIDIA | – | – | 128K |
| Free | Nemotron Nano 9B V2 (free)🧠 | NVIDIA | – | – | 128K |
| Free | gpt-oss-20b (free)🧠 | OpenAI | – | – | 131K |
| $0.08 | Ling-2.6-flash | inclusionAI | $0.01 | $0.03 | 262K |
| $0.125 | Mistral Nemo | Mistral | $0.019 | $0.03 | 131K |
| $0.197 | Granite 4.0 Micro | IBM | $0.017 | $0.112 | 131K |
| $0.225 | Nex-N2-Mini👁🧠 | Nex AGI | $0.025 | $0.1 | 262K |
| $0.25 | Llama 3 8B Lunaris | Sao10K | $0.04 | $0.05 | 8K |
| $0.28 | Qwen3.7 Flash👁🧠 | Qwen | $0.03 | $0.13 | 1M |
| $0.28 | gpt-oss-20b🧠 | OpenAI | $0.03 | $0.13 | 131K |
| $0.315 | Nova Micro 1.0 | Amazon | $0.035 | $0.14 | 128K |
| $0.325 | GPT-5 Nano (batch)👁🧠 | OpenAI | $0.025 | $0.2 | 400K |
| $0.33 | Mistral Small 3 | Mistral | $0.05 | $0.08 | 33K |
| $0.33 | Llama 3.1 8B Instruct | Meta | $0.05 | $0.08 | 131K |
| $0.336 | Llama 3.2 1B Instruct | Meta | $0.027 | $0.201 | 60K |
| $0.3375 | Command R7B (12-2024) | Cohere | $0.0375 | $0.15 | 128K |
| $0.35 | Granite 4.1 8B | IBM | $0.05 | $0.1 | 131K |
| $0.35 | Gemma 3 4B👁 | $0.05 | $0.1 | 131K | |
| $0.355 | gpt-oss-120b🧠 | OpenAI | $0.037 | $0.17 | 131K |
| $0.36 | MythoMax 13B | Other | $0.06 | $0.06 | 8K |
| $0.4 | Gemma 3 12B👁 | $0.05 | $0.15 | 131K | |
| $0.42 | Laguna XS 2.1🧠 | Poolside | $0.06 | $0.12 | 262K |
| $0.42 | Gemma 3n 4B | $0.06 | $0.12 | 33K | |
| $0.4338 | Qwen3 30B A3B Instruct 2507 | Qwen | $0.0481 | $0.193 | 262K |
| $0.45 | Nemotron 3 Nano 30B A3B🧠 | NVIDIA | $0.05 | $0.2 | 262K |
| $0.45 | Gemini 2.5 Flash Lite (batch)👁🧠 | $0.05 | $0.2 | 1.04858M | |
| $0.49 | Phi 4 | Microsoft | $0.07 | $0.14 | 16K |
| $0.525 | Hy3 preview🧠 | Tencent | $0.063 | $0.21 | 262K |
| $0.54 | Nova Lite 1.0👁 | Amazon | $0.06 | $0.24 | 300K |
| $0.58 | Llama 3.2 3B Instruct | Meta | $0.05 | $0.33 | 131K |
| $0.585 | Qwen3.5-Flash👁🧠 | Qwen | $0.065 | $0.26 | 1M |
| $0.6 | Reka Edge👁🧠 | Other | $0.1 | $0.1 | 16K |
| $0.6 | Ministral 3 3B 2512👁 | Mistral | $0.1 | $0.1 | 131K |
| $0.62 | Qwen3 Coder 30B A3B Instruct | Qwen | $0.07 | $0.27 | 262K |
| $0.65 | Qwen3.5-9B👁🧠 | Qwen | $0.1 | $0.15 | 262K |
| $0.65 | GPT-5 Nano👁🧠 | OpenAI | $0.05 | $0.4 | 400K |
| $0.675 | Seed 1.6 Flash👁🧠 | ByteDance Seed | $0.075 | $0.3 | 262K |
| $0.675 | gpt-oss-safeguard-20b🧠 | OpenAI | $0.075 | $0.3 | 131K |
| $0.68 | Qwen3 32B🧠 | Qwen | $0.08 | $0.28 | 131K |
| $0.69 | Gemma 4 26B A4B 👁🧠 | $0.07 | $0.34 | 262K | |
| $0.7 | Laguna S 2.1🧠 | Poolside | $0.1 | $0.2 | 1.04858M |
| $0.7 | GLM 4.7 Flash🧠 | Z.ai | $0.06 | $0.4 | 203K |
| $0.7 | UI-TARS 7B 👁 | ByteDance | $0.1 | $0.2 | 128K |
| $0.7 | Reka Flash 3🧠 | Other | $0.1 | $0.2 | 66K |
| $0.7 | Qwen2.5 7B Instruct | Qwen | $0.1 | $0.2 | 33K |
| $0.8 | Step 3.5 Flash🧠 | StepFun | $0.1 | $0.3 | 262K |
| $0.8 | Voxtral Small 24B 2507 | Mistral | $0.1 | $0.3 | 32K |
| $0.8 | Mistral Small 3.2 24B👁 | Mistral | $0.1 | $0.3 | 256K |
| $0.8 | Llama 4 Scout👁 | Meta | $0.1 | $0.3 | 1.31072M |
| $0.825 | Nemotron 3 Super🧠 | NVIDIA | $0.085 | $0.4 | 1M |
| $0.84 | Gemma 4 31B👁🧠 | $0.1 | $0.34 | 262K | |
| $0.85 | Gemma 3 27B👁 | $0.08 | $0.45 | 262K | |
| $0.9 | Seed-2.0-Mini👁🧠 | ByteDance Seed | $0.1 | $0.4 | 262K |
| $0.9 | Gemini 2.5 Flash Lite👁🧠 | $0.1 | $0.4 | 1.04858M | |
| $0.9 | GPT-4.1 Nano👁 | OpenAI | $0.1 | $0.4 | 1.04758M |
| $0.9 | Ministral 3 8B 2512👁 | Mistral | $0.15 | $0.15 | 262K |
| $0.936 | Qwen3 VL 32B Instruct👁 | Qwen | $0.104 | $0.416 | 131K |
| $0.98 | DeepSeek V4 Flash🧠 | DeepSeek | $0.14 | $0.28 | 1.04858M |
| $0.98 | MiMo-V2.5👁🧠 | Xiaomi | $0.14 | $0.28 | 1.05M |
| $1 | Ring-2.6-1T🧠 | inclusionAI | $0.075 | $0.625 | 262K |
| $1 | Ling-2.6-1T | inclusionAI | $0.075 | $0.625 | 262K |
| $1 | Qwen3 235B A22B Instruct 2507 | Qwen | $0.09 | $0.55 | 262K |
| $1.04 | Qwen3 VL 8B Instruct👁 | Qwen | $0.117 | $0.455 | 262K |
| $1.04 | Qwen3 8B🧠 | Qwen | $0.117 | $0.455 | 131K |
| $1.05 | Hermes 4 70B🧠 | Nous | $0.13 | $0.4 | 131K |
| $1.05 | Llama 3.3 70B Instruct | Meta | $0.13 | $0.4 | 131K |
| $1.08 | Llama Guard 4 12B👁 | Meta | $0.18 | $0.18 | 1.04858M |
| $1.1 | Qwen3 30B A3B🧠 | Qwen | $0.12 | $0.5 | 131K |
| $1.12 | GPT-5.4 Nano (batch)👁🧠 | OpenAI | $0.1 | $0.625 | 400K |
| $1.19 | Hy3🧠 | Tencent | $0.132 | $0.528 | 262K |
| $1.2 | Ministral 3 14B 2512👁 | Mistral | $0.2 | $0.2 | 262K |
| $1.25 | Olmo 3 32B Think🧠 | AllenAI | $0.15 | $0.5 | 66K |
| $1.27 | Hunyuan A13B Instruct🧠 | Tencent | $0.14 | $0.57 | 131K |
| $1.35 | KAT-Coder-Air V2.5 | Kwaipilot | $0.15 | $0.6 | 256K |
| $1.35 | MiniMax M3 (batch)👁🧠 | MiniMax | $0.15 | $0.6 | 524K |
| $1.35 | Mistral Small 4👁🧠 | Mistral | $0.15 | $0.6 | 262K |
| $1.35 | Solar Pro 3🧠 | Upstage | $0.15 | $0.6 | 128K |
| $1.35 | Qwen3 VL 30B A3B Instruct👁 | Qwen | $0.15 | $0.6 | 262K |
| $1.35 | Command R (08-2024) | Cohere | $0.15 | $0.6 | 128K |
| $1.35 | GPT-4o-mini👁 | OpenAI | $0.15 | $0.6 | 128K |
| $1.35 | GPT-4o-mini (2024-07-18)👁 | OpenAI | $0.15 | $0.6 | 128K |
| $1.38 | Gemini 3.1 Flash Lite (batch)👁🧠 | $0.125 | $0.75 | 1.04858M | |
| $1.4 | Qwen3 Coder Next | Qwen | $0.12 | $0.8 | 262K |
| $1.5 | GLM 4.5 Air🧠 | Z.ai | $0.13 | $0.85 | 131K |
| $1.6 | Saba | Mistral | $0.2 | $0.6 | 33K |
| $1.6 | Qwen3 Next 80B A3B Instruct | Qwen | $0.1 | $1.1 | 262K |
| $1.62 | GPT-5 Mini (batch)👁🧠 | OpenAI | $0.125 | $1 | 400K |
| $1.65 | MiniMax M2.5🧠 | MiniMax | $0.15 | $0.9 | 205K |
| $1.7 | Qwen3.6 35B A3B👁🧠 | Qwen | $0.14 | $1 | 262K |
| $1.7 | Qwen3.5-35B-A3B👁🧠 | Qwen | $0.14 | $1 | 262K |
| $1.74 | DeepSeek V3.2🧠 | DeepSeek | $0.269 | $0.4 | 164K |
| $1.75 | Rocinante 12B | TheDrummer | $0.25 | $0.5 | 66K |
| $1.76 | DeepSeek V3.2 Exp🧠 | DeepSeek | $0.27 | $0.41 | 164K |
| $1.8 | Llama 4 Maverick👁 | Meta | $0.2 | $0.8 | 1.04858M |
| $1.9 | Uncensored | Venice | $0.2 | $0.9 | 128K |
| $1.95 | Qwen3 Next 80B A3B Thinking🧠 | Qwen | $0.15 | $1.2 | 262K |
| $1.95 | Trinity Large Thinking🧠 | Arcee AI | $0.22 | $0.85 | 262K |
| $1.95 | Qwen3 Coder Flash | Qwen | $0.195 | $0.975 | 1M |
| $2 | Gemini 3.5 Flash Lite (batch)👁🧠 | $0.15 | $1.25 | 1.04858M | |
| $2 | Mercury 2🧠 | Inception | $0.25 | $0.75 | 128K |
| $2 | Cydonia 24B V4.1 | TheDrummer | $0.3 | $0.5 | 131K |
| $2 | Gemini 2.5 Flash (batch)👁🧠 | $0.15 | $1.25 | 1.04858M | |
| $2.05 | Qwen3 14B🧠 | Qwen | $0.2275 | $0.91 | 131K |
| $2.06 | Qwen3.6 Flash👁🧠 | Qwen | $0.1875 | $1.12 | 1M |
| $2.08 | Qwen Plus 0728🧠 | Qwen | $0.26 | $0.78 | 1M |
| $2.08 | Qwen-Plus | Qwen | $0.26 | $0.78 | 1M |
| $2.1 | MiniMax-01👁 | MiniMax | $0.2 | $1.1 | 1.00019M |
| $2.15 | Step 3.7 Flash👁🧠 | StepFun | $0.2 | $1.15 | 262K |
| $2.2 | Qwen2.5 72B Instruct | Other | $0.36 | $0.4 | 33K |
| $2.2 | DeepSeek V3.1🧠 | DeepSeek | $0.25 | $0.95 | 164K |
| $2.25 | Nex-N2-Pro👁🧠 | Nex AGI | $0.25 | $1 | 262K |
| $2.25 | Perceptron Mk1👁🧠 | Perceptron | $0.15 | $1.5 | 33K |
| $2.25 | MiniMax M2.7🧠 | MiniMax | $0.25 | $1 | 205K |
| $2.25 | GPT-5.4 Nano👁🧠 | OpenAI | $0.2 | $1.25 | 400K |
| $2.29 | MiniMax M2🧠 | MiniMax | $0.255 | $1.02 | 205K |
| $2.31 | Mistral Small 3.1 24B👁 | Mistral | $0.351 | $0.555 | 128K |
| $2.32 | DeepSeek V3 | DeepSeek | $0.2574 | $1.03 | 164K |
| $2.35 | DeepSeek V3.1 Terminus🧠 | DeepSeek | $0.27 | $1 | 164K |
| $2.4 | GLM 4.6V👁🧠 | Z.ai | $0.3 | $0.9 | 131K |
| $2.4 | Codestral 2508 | Mistral | $0.3 | $0.9 | 256K |
| $2.4 | UnslopNemo 12B | TheDrummer | $0.4 | $0.4 | 33K |
| $2.4 | Llama 3.1 70B Instruct | Meta | $0.4 | $0.4 | 131K |
| $2.47 | DeepSeek V3 0324 | DeepSeek | $0.27 | $1.12 | 164K |
| $2.5 | Qwen3 Coder 480B A35B | Qwen | $0.3 | $1 | 262K |
| $2.5 | Claude 3 Haiku👁 | Anthropic | $0.25 | $1.25 | 200K |
| $2.54 | Qwen3.5-27B👁🧠 | Qwen | $0.195 | $1.56 | 262K |
| $2.7 | LongCat 2.0🧠 | Meituan | $0.3 | $1.2 | 1.04876M |
| $2.7 | MiniMax M3👁🧠 | MiniMax | $0.3 | $1.2 | 1.04858M |
| $2.7 | KAT-Coder-Pro V2 | Kwaipilot | $0.3 | $1.2 | 262K |
| $2.7 | MiniMax M2-her | MiniMax | $0.3 | $1.2 | 66K |
| $2.7 | MiniMax M2.1🧠 | MiniMax | $0.3 | $1.2 | 205K |
| $2.75 | Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)👁🧠 | $0.25 | $1.5 | 66K | |
| $2.75 | Gemini 3.1 Flash Lite👁🧠 | $0.25 | $1.5 | 1.04858M | |
| $2.75 | Gemini 3.1 Flash Lite Preview👁🧠 | $0.25 | $1.5 | 1.04858M | |
| $2.75 | Gemini 3 Flash Preview (batch)👁🧠 | $0.25 | $1.5 | 1.04858M | |
| $2.86 | Qwen3.5 Plus 2026-02-15👁🧠 | Qwen | $0.26 | $1.56 | 1M |
| $2.88 | Qwen3.7 Plus👁🧠 | Qwen | $0.32 | $1.28 | 1M |
| $2.9 | ReMM SLERP 13B | Other | $0.45 | $0.65 | 6K |
| $2.95 | Qwen3 VL 235B A22B Instruct👁 | Qwen | $0.21 | $1.9 | 262K |
| $3 | Qwen3 VL 8B Thinking👁🧠 | Qwen | $0.18 | $2.1 | 131K |
| $3.04 | DeepSeek V4 Pro🧠 | DeepSeek | $0.435 | $0.87 | 1.04858M |
| $3.04 | MiMo-V2.5-Pro🧠 | Xiaomi | $0.435 | $0.87 | 1.05M |
| $3.2 | Qwen Plus 0728 (thinking)🧠 | Qwen | $0.4 | $1.2 | 1M |
| $3.25 | Seed-2.0-Lite👁🧠 | ByteDance Seed | $0.25 | $2 | 262K |
| $3.25 | Seed 1.6👁🧠 | ByteDance Seed | $0.25 | $2 | 262K |
| $3.25 | GPT-5.1-Codex-Mini👁🧠 | OpenAI | $0.25 | $2 | 400K |
Common questions
What's the cheapest LLM API?
Proprietary: Gemini Flash-Lite and GPT-4.1 Nano near $0.10/1M input. Cheapest overall is DeepSeek (~$0.14 in / $0.28 out). The best pick depends on your input:output ratio, use the calculator.
Why do input and output cost different amounts?
Generating output is far more expensive than reading input, so output is usually 3-6× the input price. That's why a single blended price is misleading.
Is Claude or GPT cheaper?
Close at the frontier. Claude Opus ~$5/$25, GPT-5 ~$2.50/$15 standard (much more for Pro). Claude's prompt caching can cut input up to 90%, often making it cheaper in practice.
How do I cut LLM costs?
Use a smaller model where you can (budget models are 20-50× cheaper), cache repeated prompts, use the batch API for non-urgent jobs, and trim prompt and output length.