AI API PricingCost per Token 2026

How much does it cost to use GPT, Claude, Gemini and other LLM APIs? Compare pricing per million tokens — input and output rates for every major model.

Last updated: August 31, 2026 719 APIs listed

719

APIs listed

193

With free tier

$0.019

Cheapest (input/1M)

$150.00

Most expensive (input/1M)

How to read the table: prices are per million tokens (input = what you send; output = the model's response). In English, 1,000 tokens ≈ 750 words ≈ 1 A4 page. Prices verified on each company's official pricing page.

Price per Million Tokens — Paid APIs

#ModelInput USD/1M
1Mistral: Mistral Nemo$0.019
2Gemma 4 E4B (Non-reasoning)$0.020
3Gemma 4 E4B (Reasoning)$0.020
4Llama 3.1 8B Instruct$0.020
5OpenAI: GPT-5 Nano (batch)$0.025
6Qwen3.5 4B (Non-reasoning)$0.030
7Qwen3.5 4B (Reasoning)$0.030
8Granite 3.3 8B (Non-reasoning)$0.030
9Granite 4.2 3B$0.030
10Sarvam 30B$0.030
11Qwen: Qwen2.5 7B Instruct$0.040
12Amazon: Nova Micro 1.0$0.040
13Llama 3 8B Instruct$0.040
14HyperNova 60B 2605$0.040
15NVIDIA Nemotron Nano 9B V2 (Reasoning)$0.040
16Sao10K: Llama 3 8B Lunaris$0.040
17Sarvam 105B (high)$0.040
18Arcee AI: Trinity Mini$0.045
19Qwen: Qwen-Turbo$0.050
20Google: Gemini 2.5 Flash Lite (batch)$0.050
21Granite 4.1 8B$0.050
22Llama 2 Chat 7B$0.050
23NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)$0.050
24NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)$0.050
25NVIDIA Nemotron Nano 9B V2 (Non-reasoning)$0.050
26NVIDIA: Nemotron 3 Nano 30B A3B$0.050
27NVIDIA: Nemotron Nano 9B V2$0.050
28GPT-5 Nano$0.050
29GPT-5 nano (minimal)$0.050
30OpenAI: GPT-4.1 Nano (batch)$0.050
31OpenAI: gpt-oss-20b (batch)$0.050
32Amazon: Nova Lite 1.0$0.060
33Gemma 3n 4B$0.060
34Gemma 3n E4B Instruct$0.060
35Granite 4.0 H Small$0.060
36Granite 4.2 8B$0.060
37MythoMax 13B$0.060
38gpt-oss-20b$0.060
39Hy3-preview (Non-reasoning)$0.060
40Ling-3.0-flash$0.070
41Baidu: ERNIE 4.5 21B A3B Thinking$0.070
42Nemotron 3 Nano Omni 30B A3B Reasoning$0.070
43Nemotron 3.5 Lightning$0.070
44Z.ai: GLM 4.7 Flash$0.070
45ByteDance Seed: Seed 1.6 Flash$0.075
46Gemini 2.0 Flash Lite$0.075
47gpt-oss-safeguard-20b$0.075
48OpenAI: GPT-4o-mini (batch)$0.075
49Qwen: Qwen3 30B A3B Thinking 2507$0.080
50DeepSeek: DeepSeek V4 Flash 0423$0.089
51Tongyi DeepResearch 30B A3B$0.090
52Qwen3.5 Omni Flash$0.100
53Olmo 3 7B Instruct$0.100
54ByteDance: UI-TARS 7B $0.100
55Gemini 2.5 Flash Lite$0.100
56Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning)$0.100
57Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning)$0.100
58Gemma 4 12B (Reasoning)$0.100
59Ling 2.6 Flash$0.100
60Ministral 3 3B$0.100
61Mistral Small 3$0.100
62Mistral Small 3.1$0.100
63Mistral Small 3.2$0.100
64Mistral: Voxtral Small 24B 2507$0.100
65GPT-4.1 Nano$0.100
66OpenAI: GPT-5.4 Nano (batch)$0.100
67OpenAI: GPT-5.6 Luna (batch)$0.100
68OpenAI: GPT-5.6 Luna Pro (batch)$0.100
69Reka Edge$0.100
70Agnes 2.5 Pro Beta$0.100
71Step 3.5 Flash$0.100
72Apertus 8B Instruct$0.100
73MiMo-V2-Flash (Reasoning)$0.100
74Z.ai: GLM 4 32B $0.100
75Mistral: Ministral 8B$0.110
76Google: Gemini 3.1 Flash Lite (batch)$0.125
77OpenAI: GPT-5 Mini (batch)$0.125
78DeepSeek V4 Flash (Reasoning, High Effort)$0.130
79DeepSeek V4 Flash (Reasoning, Max Effort)$0.130
80Gemma 4 26B A4B $0.130
81Microsoft: Phi 4$0.130
82Nous: Hermes 4 70B$0.130
83Hermes 4 - Llama-3.1 70B (Non-reasoning)$0.130
84Hermes 4 - Llama-3.1 70B (Reasoning)$0.130
85Nex AGI: DeepSeek V3.1 Nex N1$0.135
86Qwen3.5 9B (Reasoning)$0.140
87Baidu: ERNIE 4.5 VL 28B A3B$0.140
88DeepSeek-V4-Flash$0.140
89DeepSeek: DeepSeek V4 Flash 0731 (batch)$0.140
90Ling-flash-2.0$0.140
91Ring-flash-2.0$0.140
92NousResearch: Hermes 2 Pro - Llama-3 8B$0.140
93Hy3-preview (Reasoning)$0.140
94Tencent: Hunyuan A13B Instruct$0.140
95MiMo-V2.5$0.140
96Qwen: Qwen3 235B A22B Thinking 2507$0.149
97Qwen: Qwen3 Next 80B A3B Instruct$0.150
98Qwen3 Next 80B A3B (Reasoning)$0.150
99Qwen3.8-Flash-Next$0.150
100EssentialAI: Rnj 1 Instruct$0.150
101Google: Gemini 2.5 Flash (batch)$0.150
102Google: Gemini 3.5 Flash Lite (batch)$0.150
103Ministral 3 8B$0.150
104Mistral Small (Feb '24)$0.150
105Mistral: Ministral 3 8B 2512 (batch)$0.150
106Mistral: Mistral Small 4$0.150
107Mistral: Mistral Small 4 (batch)$0.150
108GPT-4o-mini (2024-07-18)$0.150
109GPT-4o-mini Search Preview$0.150
110gpt-oss-120b$0.150
111OpenAI: GPT-4o-mini$0.150
112OpenAI: gpt-oss-120b (batch)$0.150
113Solar Mini$0.150
114Solar Pro 3$0.150
115GLM-5.3-Flash$0.150
116Qwen: Qwen3 VL 32B Instruct$0.160
117Qwen3 32B (Non-reasoning)$0.160
118Qwen3 32B (Reasoning)$0.160
119Qwen3 VL 32B (Reasoning)$0.160
120Granite 4.2 30B$0.160
121TheDrummer: Rocinante 12B$0.170
122Z.ai: GLM 4.5 Air$0.170
123Qwen: Qwen3 VL 8B Instruct$0.180
124Qwen3 8B (Non-reasoning)$0.180
125Qwen3 8B (Reasoning)$0.180
126Qwen3 VL 8B (Reasoning)$0.180
127Arcee AI: Spotlight$0.180
128Llama 4 Scout$0.180
129Llama Guard 4 12B$0.180
130Google: Gemini 3.7 Flash (batch)$0.188
131Jamba 1.5 Mini$0.200
132Jamba 1.6 Mini$0.200
133Qwen: Qwen3 30B A3B Instruct 2507$0.200
134Qwen: Qwen3 VL 30B A3B Instruct$0.200
135Qwen3 30B A3B (Reasoning)$0.200
136Qwen3 30B A3B 2507 (Reasoning)$0.200
137Qwen3 30B A3B 2507 Instruct$0.200
138Qwen3 VL 30B A3B (Reasoning)$0.200
139MiniMax: MiniMax-01$0.200
140Ministral 3 14B$0.200
141Mistral Small (Sep '24)$0.200
142Mistral: Saba$0.200
143NVIDIA Nemotron 3 Super 120B A12B (Reasoning)$0.200
144NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)$0.200
145NVIDIA Nemotron Nano 12B v2 VL (Reasoning)$0.200
146GPT-5.4 Nano$0.200
147GPT-5.6 Luna (high)$0.200
148GPT-5.6 Luna (low)$0.200
149GPT-5.6 Luna (max)$0.200
150GPT-5.6 Luna (medium)$0.200
151GPT-5.6 Luna (Non-reasoning)$0.200
152GPT-5.6 Luna (xhigh)$0.200
153OpenAI: GPT-4.1 Mini (batch)$0.200
154OpenAI: GPT-5.6 Luna Pro$0.200
155Reka Flash 3$0.200
156Step 3.7 Flash$0.200
157Celeris-1$0.200
158Grok 4 Fast$0.200
159Seed-OSS-36B-Instruct$0.210
160Qwen: Qwen3 235B A22B Instruct 2507$0.230
161Arcee AI: Trinity Large Thinking$0.230
162Qwen: Qwen2.5 VL 72B Instruct$0.250
163Qwen: Qwen3.5-35B-A3B$0.250
164Qwen3 Omni 30B A3B (Reasoning)$0.250
165Qwen3 Omni 30B A3B Instruct$0.250
166Anthropic: Claude 3 Haiku$0.250
167ByteDance Seed: Seed-2.0-Lite$0.250
168Gemini 3.1 Flash Lite$0.250
169Gemini 3.1 Flash Lite Preview$0.250
170Google: Gemini 3 Flash Preview (batch)$0.250
171Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.250
172Inception: Mercury 2$0.250
173Mistral 7B Instruct$0.250
174GPT-5 Mini$0.250
175GPT-5 mini (minimal)$0.250
176GPT-5.1-Codex-Mini$0.250
177OpenAI: GPT-3.5 Turbo (batch)$0.250
178Llama 4 Maverick$0.260
179DeepSeek V3 0324$0.270
180DeepSeek V3.1 Terminus$0.270
181DeepSeek V3.2 Exp$0.270
182Baidu: ERNIE 4.5 300B A47B $0.280
183DeepSeek V3.2$0.280
184DeepSeek V3.2 Exp (Non-reasoning)$0.280
185DeepSeek V3.2 Exp (Reasoning)$0.280
186Qwen: Qwen3.5-27B$0.300
187Amazon: Nova 2 Lite$0.300
188Nova 2.0 Lite (high)$0.300
189Nova 2.0 Omni (low)$0.300
190Nova 2.0 Omni (medium)$0.300
191Nova 2.0 Omni (Non-reasoning)$0.300
192Apodex 1.1$0.300
193Gemini 2.5 Flash$0.300
194Gemini 2.5 Flash Preview (Reasoning)$0.300
195Gemini 3.5 Flash-Lite$0.300
196Nano Banana (Gemini 2.5 Flash Image)$0.300
197Ling-2.6-1T$0.300
198Ring-2.6-1T$0.300
199Kwaipilot: KAT-Coder-Pro V2$0.300
200MiniMax-M2$0.300
201MiniMax-M3$0.300
202MiniMax: MiniMax M2-her$0.300
203MiniMax: MiniMax M2.1$0.300
204MiniMax: MiniMax M2.5$0.300
205MiniMax: MiniMax M2.7$0.300
206Inkling-Small$0.300
207Mistral: Codestral 2508$0.300
208Mistral: Codestral 2508 (batch)$0.300
209Nous: Hermes 3 70B Instruct$0.300
210TheDrummer: Cydonia 24B V4.1$0.300
211Inkling Small$0.300
212Solar Pro 4$0.300
213Grok 3 Mini$0.300
214Grok 3 Mini Beta$0.300
215GLM-4.6V (Reasoning)$0.300
216Muse Glimmer (high)$0.320
217Llama 3.2 11B Vision Instruct$0.340
218Qwen: Qwen3 Coder Next$0.350
219Qwen3 14B (Non-reasoning)$0.350
220Qwen3 14B (Reasoning)$0.350
221DeepSeek V3$0.360
222Google: Gemini 3.6 Flash (batch)$0.375
223OpenAI: GPT-5.4 Mini (batch)$0.375
224Qwen: Qwen3.6 35B A3B$0.380
225Google: Gemma 4 31B (batch)$0.390
226Qwen: Qwen3 VL 235B A22B Instruct$0.400
227Qwen: Qwen3.5-122B-A10B$0.400
228Qwen3 VL 235B A22B (Reasoning)$0.400
229Qwen3.5 Omni Plus$0.400
230Qwen3.7 Plus$0.400
231MiniMax: MiniMax M1$0.400
232Mistral: Mistral Medium 3$0.400
233Mistral: Mistral Medium 3.1$0.400
234Mistral: Mistral Medium 3.1 (batch)$0.400
235Llama Nemotron Super 49B v1.5 (Non-reasoning)$0.400
236Llama Nemotron Super 49B v1.5 (Reasoning)$0.400
237GPT-4.1 Mini$0.400
238TheDrummer: UnslopNemo 12B$0.400
239Baidu: ERNIE 4.5 VL 424B A47B $0.420
240DeepSeek V4 Pro (Non-reasoning)$0.430
241DeepSeek V4 Pro (Reasoning, High Effort)$0.430
242DeepSeek V4 Pro (Reasoning, Max Effort)$0.430
243Xiaomi: MiMo-V2.5-Pro$0.430
244DeepSeek V4 Flash$0.440
245DeepSeek V4 Flash Vision (Reasoning, Max Effort)$0.440
246DeepSeek: DeepSeek V4 Flash Vision Exp$0.440
247Qwen: Qwen3 Coder 30B A3B Instruct$0.450
248Mistral: Mixtral 8x7B Instruct$0.450
249ReMM SLERP 13B$0.450
250Agnes 2.5 Pro Alpha$0.450
251Qwen2.5 72B Instruct$0.470
252Llama Guard 3 8B$0.480
253Qwen: Qwen3.6 Plus$0.500
254Qwen: Qwen3.8 27B$0.500
255Anthropic: Claude Haiku 4.5 (batch)$0.500
256Arcee AI: Coder Large$0.500
257Command-R (Mar '24)$0.500
258Gemini 3 Flash Preview (Non-reasoning)$0.500
259Gemini 3 Flash Preview (Reasoning)$0.500
260Google: Nano Banana 2 (Gemini 3.1 Flash Image)$0.500
261Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.500
262Magistral Small 1.2$0.500
263Mistral Large 3$0.500
264Mistral: Mistral Large 3 2512 (batch)$0.500
265Nex-N2-Pro$0.500
266GPT-3.5 Turbo$0.500
267MiniMax M1 80k$0.550
268OpenAI: o3 Mini (batch)$0.550
269OpenAI: o3 Mini High (batch)$0.550
270OpenAI: o4 Mini (batch)$0.550
271OpenAI: o4 Mini High (batch)$0.550
272TheDrummer: Skyfall 36B V2$0.550
273Llama 3.1 70B Instruct$0.560
274DeepSeek V3.1$0.570
275Kimi K2$0.570
276MoonshotAI: Kimi K2 0711$0.570
277GLM-4.6 (Reasoning)$0.570
278Qwen: Qwen3.5 397B A17B$0.600
279Qwen: Qwen3.6 27B$0.600
280Kimi K2 Thinking$0.600
281MoonshotAI: Kimi K2 0905$0.600
282MoonshotAI: Kimi K2.5$0.600
283Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)$0.600
284GPT Audio Mini$0.600
285Writer: Palmyra X5$0.600
286GLM-4.5V (Reasoning)$0.600
287GLM-4.7 (Reasoning)$0.600
288WizardLM-2 8x22B$0.620
289Google: Gemini 2.5 Pro (batch)$0.625
290OpenAI: GPT-5 (batch)$0.625
291OpenAI: GPT-5 Codex (batch)$0.625
292OpenAI: GPT-5.1 (batch)$0.625
293Gemma 2 27B$0.650
294Llama 3 70B Instruct$0.650
295Sao10K: Llama 3.3 Euryale 70B$0.650
296QwQ 32B$0.660
297Llama 3.3 70B Instruct$0.660
298Nemotron 3 Ultra 550B A55B (Reasoning)$0.680
299Qwen3 235B A22B (Reasoning)$0.700
300DeepSeek: R1$0.700
301R1 Distill Llama 70B$0.700
302Hermes 3 - Llama-3.1 70B$0.700
303Arcee AI: Virtuoso Large$0.750
304Gemini 3.6 Flash (high)$0.750
305Gemini 3.7 Flash (high)$0.750
306Gemini 3.7 Flash (low)$0.750
307Gemini 3.7 Flash (medium)$0.750
308Google: Gemini 3.5 Flash (batch)$0.750
309LongCat 2.0$0.750
310Mancer: Weaver (alpha)$0.750
311Mistral: Mistral Medium 3.5 (batch)$0.750
312GPT-5.4 Mini$0.750
313Qwen: Qwen3 Max Thinking$0.780
314AionLabs: Aion-2.0$0.800
315AionLabs: Aion-RP 1.0 (8B)$0.800
316AlfredPros: CodeLLaMa 7B Instruct Solidity$0.800
317Amazon: Nova Pro 1.0$0.800
318Morph: Morph V3 Fast$0.800
319Apertus 70B Instruct$0.820
320Relace: Relace Apply 3$0.850
321Sao10K: Llama 3.1 Euryale 70B v2.2$0.850
322OpenAI: GPT-5.2 (batch)$0.875
323Arcee AI: Maestro Reasoning$0.900
324Morph: Morph V3 Large$0.900
325Kimi K2.7 Code$0.950
326MoonshotAI: Kimi K2.6$0.950
327Anthropic: Claude Sonnet 5 (batch)$1.00
328Claude 4.5 Haiku (Reasoning)$1.00
329Claude Haiku 4.5$1.00
330Google: Gemini 3.1 Pro Preview (batch)$1.00
331Nous: Hermes 3 405B Instruct$1.00
332Nous: Hermes 4 405B$1.00
333Hermes 4 - Llama-3.1 405B (Non-reasoning)$1.00
334Hermes 4 - Llama-3.1 405B (Reasoning)$1.00
335OpenAI: GPT-4.1 (batch)$1.00
336OpenAI: GPT-5.6 Sol (batch)$1.00
337OpenAI: GPT-5.6 Sol Pro (batch)$1.00
338OpenAI: GPT-5.6 Terra (batch)$1.00
339OpenAI: GPT-5.6 Terra Pro (batch)$1.00
340OpenAI: o3 (batch)$1.00
341Relace: Relace Search$1.00
342Inkling$1.00
343Inkling$1.00
344Grok Build 0.1 0616$1.00
345SpaceXAI: Grok Build 0.1$1.00
346xAI: Grok Build 0.1$1.00
347GLM-5 (Non-reasoning)$1.00
348o3 Mini$1.10
349o3 Mini High$1.10
350o4 Mini$1.10
351o4 Mini High$1.10
352Qwen: Qwen3 Max$1.20
353Qwen3 Max (Preview)$1.20
354Qwen3 Max Thinking (Preview)$1.20
355Llama 3.1 Nemotron 70B Instruct$1.20
356Z.ai: GLM 5V Turbo$1.20
357Nova 2.0 Pro Preview (medium)$1.25
358Cogito v2.1 (Reasoning)$1.25
359Deep Cogito: Cogito v2.1 671B$1.25
360Gemini 2.5 Pro$1.25
361Gemini 2.5 Pro Preview (May' 25)$1.25
362Gemini 2.5 Pro Preview 05-06$1.25
363Muse Spark 1.1$1.25
364Muse Spark 1.1 (xhigh)$1.25
365Muse Spark 1.2 (xhigh)$1.25
366GPT-5$1.25
367GPT-5 (minimal)$1.25
368GPT-5 Codex$1.25
369GPT-5.1$1.25
370GPT-5.1 Chat$1.25
371GPT-5.1-Codex$1.25
372GPT-5.1-Codex-Max$1.25
373OpenAI: GPT-4o (batch)$1.25
374OpenAI: GPT-5.4 (batch)$1.25
375Grok 4.20 Multi-Agent$1.25
376SpaceXAI: Grok 4.20$1.25
377SpaceXAI: Grok 4.20 Multi-Agent$1.25
378SpaceXAI: Grok 4.3$1.25
379Z.ai: GLM 5.1$1.28
380Qwen3.6 Max Preview$1.30
381DeepSeek V4 Pro$1.32
382DeepSeek: DeepSeek V4 Pro 0813 (batch)$1.32
383DeepSeek R1 (Jan '25)$1.35
384GLM-5.1 (Non-reasoning)$1.38
385GLM-5.2 (max)$1.40
386GLM-5.3$1.40
387Sao10k: Llama 3 Euryale 70B v2.1$1.48
388Qwen3 Coder 480B A35B Instruct$1.50
389Anthropic: Claude Sonnet 4.5 (batch)$1.50
390Anthropic: Claude Sonnet 4.6 (batch)$1.50
391Gemini 3.5 Flash$1.50
392Gemini 3.5 Flash (minimal)$1.50
393Gemini 3.6 Flash$1.50
394Google: Gemini 3.5 Flash$1.50
395Mistral Medium$1.50
396Mistral: Mistral Medium 3.5$1.50
397DeepSeek: DeepSeek V4 Pro 0423$1.60
398GPT-5.2$1.75
399GPT-5.2-Codex$1.75
400GPT-5.3 Chat$1.75
401GPT-5.3-Codex$1.75
402AI21: Jamba Large 1.7$2.00
403Jamba 1.5 Large$2.00
404Jamba 1.6 Large$2.00
405Qwen: Qwen3.8 2.4T A95B$2.00
406Qwen: Qwen3.8 Max$2.00
407Claude Sonnet 5$2.00
408Gemini 3 Pro Preview (high)$2.00
409Gemini 3 Pro Preview (low)$2.00
410Gemini 3.1 Pro Preview$2.00
411Gemini 3.1 Pro Preview Custom Tools$2.00
412Google: Nano Banana Pro (Gemini 3 Pro Image)$2.00
413Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00
414Mistral Large 2 (Jul '24)$2.00
415Magistral Medium 1.2$2.00
416Mistral Large$2.00
417Mistral: Mixtral 8x22B Instruct$2.00
418GPT-4.1$2.00
419GPT-5.6 Terra (high)$2.00
420GPT-5.6 Terra (low)$2.00
421GPT-5.6 Terra (max)$2.00
422GPT-5.6 Terra (medium)$2.00
423GPT-5.6 Terra (Non-reasoning)$2.00
424GPT-5.6 Terra (xhigh)$2.00
425o3$2.00
426o4 Mini Deep Research$2.00
427OpenAI: GPT-5.6 Sol Pro$2.00
428OpenAI: GPT-5.6 Terra Pro$2.00
429Perplexity: Sonar Deep Research$2.00
430Grok 4.5$2.00
431Grok 4.20 0309 (Reasoning)$2.00
432SpaceXAI: Grok 4.5$2.00
433SpaceXAI: Grok 4.6$2.00
434Qwen3.7 Max$2.50
435Nova Premier$2.50
436Anthropic: Claude Opus 4.5 (batch)$2.50
437Anthropic: Claude Opus 4.6 (batch)$2.50
438Anthropic: Claude Opus 4.7 (batch)$2.50
439Anthropic: Claude Opus 4.8 (batch)$2.50
440Claude Opus 5 (batch)$2.50
441Cohere: Command A$2.50
442Cohere: Command R+ (08-2024)$2.50
443Inflection: Inflection 3 Pi$2.50
444Inflection: Inflection 3 Productivity$2.50
445Llama 3.1 Instruct 405B$2.50
446GPT Audio$2.50
447GPT-4o (2024-08-06)$2.50
448GPT-4o (2024-11-20)$2.50
449GPT-4o Audio$2.50
450GPT-4o Search Preview$2.50
451GPT-5 Image Mini$2.50
452GPT-5.4$2.50
453OpenAI: GPT-4o$2.50
454OpenAI: GPT-5.5 (batch)$2.50
455Claude 3 Sonnet$3.00
456Claude 3.5 Sonnet (June '24)$3.00
457Claude 3.5 Sonnet (Oct '24)$3.00
458Claude 3.7 Sonnet$3.00
459Claude 4 Sonnet (Reasoning)$3.00
460Claude 4.5 Sonnet (Non-reasoning)$3.00
461Claude 4.5 Sonnet (Reasoning)$3.00
462Claude Sonnet 4$3.00
463Claude Sonnet 4.5$3.00
464Claude Sonnet 4.6$3.00
465Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)$3.00
466Claude Sonnet 4.6 (Non-reasoning, Low Effort)$3.00
467Command-R+ (Apr '24)$3.00
468Kimi K3$3.00
469Magnum v4 72B$3.00
470OpenAI: GPT-3.5 Turbo 16k$3.00
471Perplexity: Sonar Pro Search$3.00
472Sao10K: Llama 3.1 70B Hanami x1$3.00
473Grok 3 Beta$3.00
474Grok 4$3.00
475Goliath 120B$3.75
476AionLabs: Aion-1.0$4.00
477Mistral Large 2 (Nov '24)$4.00
478GPT-5.6 Sol (high)$4.00
479GPT-5.6 Sol (low)$4.00
480GPT-5.6 Sol (max)$4.00
481GPT-5.6 Sol (medium)$4.00
482GPT-5.6 Sol (Non-reasoning)$4.00
483GPT-5.6 Sol (xhigh)$4.00
484Grok 3$4.00
485Anthropic: Claude Fable 5 (batch)$5.00
486Claude Opus 4.5$5.00
487Claude Opus 4.5 (Reasoning)$5.00
488Claude Opus 4.6$5.00
489Claude Opus 4.6 (Adaptive Reasoning, Max Effort)$5.00
490Claude Opus 4.7$5.00
491Claude Opus 4.8 (Adaptive Reasoning, Max Effort)$5.00
492Claude Opus 5$5.00
493GPT Chat Latest$5.00
494GPT-5.5$5.00
495GPT-5.5 Instant (June 2026)$5.00
496GPT-5.5 Instant (May 2026)$5.00
497OpenAI: GPT-4 Turbo (batch)$5.00
498OpenAI: GPT-4o (2024-05-13)$5.00
499Anthropic: Claude Opus 4.1 (batch)$7.50
500OpenAI: GPT-5 Pro (batch)$7.50
501OpenAI: o1 (batch)$7.50
502GPT-5.4 Image 2$8.00
503Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)$10.00
504Claude Opus 5 (Fast)$10.00
505GPT-4 Turbo$10.00
506GPT-5 Image$10.00
507o3 Deep Research$10.00
508OpenAI: GPT-4 Turbo (older v1106)$10.00
509OpenAI: o3 Pro (batch)$10.00
510OpenAI: GPT-5.2 Pro (batch)$10.50
511Claude 3 Opus$15.00
512Claude 4 Opus (Reasoning)$15.00
513Claude 4.1 Opus (Non-reasoning)$15.00
514Claude 4.1 Opus (Reasoning)$15.00
515Claude Opus 4$15.00
516Claude Opus 4.1$15.00
517o1$15.00
518OpenAI: GPT-5.4 Pro (batch)$15.00
519OpenAI: GPT-5.5 Pro (batch)$15.00
520o1-preview$16.50
521o3 Pro$20.00
522Claude Opus 4.7 (Fast)$30.00
523GPT-5.4 Pro$30.00
524OpenAI: GPT-4$30.00
525OpenAI: o1-pro (batch)$75.00
526o1-pro$150.00

APIs with Free Tier

These models offer free API access (with rate limits). Ideal for prototypes and low-volume projects.

Apriel-v1.5-15B-Thinker

ServiceNow

Free

Apriel-v1.6-15B-Thinker

ServiceNow

Free

Arctic Instruct

Snowflake

Free

Claude 2.0

Anthropic

Free

Claude 2.1

Anthropic

Free

Claude 3.5 Haiku

Anthropic

Free

Claude 3.7 Sonnet (thinking)

Anthropic

Free

Claude Instant

Anthropic

Free

Command A+

Cohere

Free

DBRX Instruct

Databricks

Free

DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)

Nous Research

Free

DeepHermes 3 - Mistral 24B Preview (Non-reasoning)

Nous Research

Free

DeepSeek Coder V2 Lite Instruct

DeepSeek

Free

DeepSeek LLM 67B Chat (V1)

DeepSeek

Free

DeepSeek R1 0528 Qwen3 8B

DeepSeek

Free

DeepSeek R1 Distill Llama 8B

DeepSeek

Free

DeepSeek R1 Distill Qwen 1.5B

DeepSeek

Free

DeepSeek R1 Distill Qwen 14B

DeepSeek

Free

DeepSeek V3.2 Speciale

DeepSeek

Free

DeepSeek-Coder-V2

DeepSeek

Free

DeepSeek-V2-Chat

DeepSeek

Free

DeepSeek-V2.5

DeepSeek

Free

DeepSeek-V2.5 (Dec '24)

DeepSeek

Free

DeepSeek: R1 Distill Qwen 32B

DeepSeek

Free

Devstral 2

Mistral

Free

Devstral Small (Jul '25)

Mistral

Free

Devstral Small (May '25)

Mistral

Free

Devstral Small 2

Mistral

Free

DiffusionGemma 26B A4B

Google

Free

Doubao Seed Code

ByteDance

Free

ERNIE 5.0 Thinking Preview

Baidu

Free

Exaone 4.0 1.2B (Non-reasoning)

LG AI Research

Free

EXAONE 4.0 32B (Non-reasoning)

LG AI Research

Free

EXAONE 4.0 32B (Reasoning)

LG AI Research

Free

EXAONE 4.5 33B

LG AI

Free

Falcon-H1R-7B

TII UAE

Free

G9v3-39A5B

AI9Stars

Free

G9v3-3B

AI9Stars

Free

Gemini 1.0 Pro

Google

Free

Gemini 1.0 Ultra

Google

Free

Gemini 1.5 Flash (May '24)

Google

Free

Gemini 1.5 Flash (Sep '24)

Google

Free

Gemini 1.5 Flash-8B

Google

Free

Gemini 1.5 Pro (May '24)

Google

Free

Gemini 1.5 Pro (Sep '24)

Google

Free

Gemini 2.0 Flash

Google

Free

Gemini 2.0 Flash (experimental)

Google

Free

Gemini 2.0 Flash Thinking Experimental (Dec '24)

Google

Free

Gemini 2.0 Flash Thinking Experimental (Jan '25)

Google

Free

Gemini 2.0 Flash-Lite (Feb '25)

Google

Free

Gemini 2.0 Flash-Lite (Preview)

Google

Free

Gemini 2.0 Pro Experimental (Feb '25)

Google

Free

Gemini 2.5 Flash Preview (Non-reasoning)

Google

Free

Gemini 2.5 Flash Preview (Sep '25) (Reasoning)

Google

Free

Gemini 2.5 Pro Preview (Mar' 25)

Google

Free

Gemini 3 Deep Think

Google

Free

Gemma 3 12B

Google

Free

Gemma 3 1B Instruct

Google

Free

Gemma 3 270M

Google

Free

Gemma 3 27B

Google

Free

Gemma 3 4B

Google

Free

Gemma 3n E2B Instruct

Google

Free

Gemma 3n E4B Instruct Preview (May '25)

Google

Free

Gemma 4 31B

Google

Free

Gemma 4 E2B (Non-reasoning)

Google

Free

Gemma 4 E2B (Reasoning)

Google

Free

GLM-4.5 (Reasoning)

Z.ai

Free

GLM-5-Turbo

Z.ai

Free

GPT-3.5 Turbo (0613)

OpenAI

Free

GPT-4.5 (Preview)

OpenAI

Free

GPT-4o (ChatGPT)

OpenAI

Free

GPT-4o (March 2025, chatgpt-4o-latest)

OpenAI

Free

GPT-4o mini Realtime (Dec '24)

OpenAI

Free

GPT-4o Realtime (Dec '24)

OpenAI

Free

GPT-5 (ChatGPT)

OpenAI

Free

Granite 4.0 1B

IBM

Free

Granite 4.0 350M

IBM

Free

Granite 4.0 H 1B

IBM

Free

Granite 4.0 H 350M

IBM

Free

Granite 4.0 Micro

IBM

Free

Granite 4.1 30B

IBM

Free

Granite 4.1 3B

IBM

Free

Grok 2 (Dec '24)

xAI

Free

Grok 4.1 Fast

xAI

Free

Grok Beta

xAI

Free

Grok Code Fast 1

xAI

Free

Grok-1

xAI

Free

HyperCLOVA X SEED Think (32B)

Naver

Free

INTELLECT-3

Prime Intellect

Free

Jamba 1.7 Mini

AI21 Labs

Free

Jamba Reasoning 3B

AI21 Labs

Free

JT-35B-Flash

China Mobile

Free

JT-4.1 Flash 236B A21B

China Mobile

Free

JT-MINI

China Mobile

Free

K2 Think V2

MBZUAI Institute of Foundation Models

Free

K2-V2 (medium)

MBZUAI Institute of Foundation Models

Free

KAT-Coder-Pro V1

KwaiKAT

Free

Kimi Linear 48B A3B Instruct

Kimi

Free

LFM 40B

Liquid AI

Free

LFM2 1.2B

Liquid AI

Free

LFM2 2.6B

Liquid AI

Free

LFM2 8B A1B

Liquid AI

Free

LFM2-24B-A2B

LiquidAI

Free

LFM2.5-1.2B-Instruct

Liquid AI

Free

LFM2.5-1.2B-Thinking

Liquid AI

Free

LFM2.5-2.6B

Free

LFM2.5-8B-A1B

Liquid AI

Free

LFM2.5-VL-1.6B

Liquid AI

Free

Ling 3.0 Tiny

InclusionAI

Free

Ling-1T

InclusionAI

Free

Ling-mini-2.0

InclusionAI

Free

Llama 2 Chat 13B

Meta

Free

Llama 2 Chat 70B

Meta

Free

Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)

Nvidia

Free

Llama 3.1 Tulu3 405B

Allen Institute for AI

Free

Llama 3.2 1B Instruct

Meta

Free

Llama 3.2 3B Instruct

Meta

Free

Llama 3.2 Instruct 90B (Vision)

Meta

Free

Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)

Nvidia

Free

Llama 3.3 Nemotron Super 49B v1 (Reasoning)

Nvidia

Free

Llama 65B

Meta

Free

LongCat Flash Lite

LongCat

Free

Magistral Medium 1

Mistral

Free

Magistral Small 1

Mistral

Free

Mi:dm K 2.5 Pro

Korea Telecom

Free

Mi:dm K 2.5 Pro Preview

Korea Telecom

Free

MiMo-V2-Flash (Feb 2026)

Xiaomi

Free

MiMo-V2-Omni-0327

Xiaomi

Free

MiniCPM-V 4.6 1.3B

OpenBMB

Free

MiniCPM5-1B (Non-reasoning)

OpenBMB

Free

MiniMax M1 40k

MiniMax

Free

Mistral: Devstral Medium

Mistral AI

Free

Mistral: Pixtral Large 2411

Mistral AI

Free

Mixtral 8x22B Instruct

Mistral

Free

Molmo 7B-D

Allen Institute for AI

Free

Molmo2-8B

Allen Institute for AI

Free

Motif 3 (Beta)

Motif Technologies

Free

Motif-2-12.7B-Reasoning

Motif Technologies

Free

Muse Spark

Meta

Free

Nanbeige4.1-3B

Nanbeige

Free

Nemotron Cascade 2 30B A3B

Nvidia

Free

North Mini Code

Cohere

Free

NVIDIA Nemotron 3 Nano 4B

Nvidia

Free

o1-mini

OpenAI

Free

OLMo 2 32B

Allen Institute for AI

Free

OLMo 2 7B

Allen Institute for AI

Free

Olmo 3 32B Think

AllenAI

Free

Olmo 3 7B Think

Allen Institute for AI

Free

Olmo 3.1 32B Instruct

AllenAI

Free

Olmo 3.1 32B Think

Allen Institute for AI

Free

OpenChat 3.5 (1210)

OpenChat

Free

PALM-2

Google

Free

Phi-3 Mini Instruct 3.8B

Microsoft

Free

Phi-4 Mini Instruct

Microsoft

Free

Phi-4 Multimodal Instruct

Microsoft

Free

Qwen Chat 14B

Alibaba

Free

Qwen Chat 72B

Alibaba

Free

Qwen1.5 Chat 110B

Alibaba

Free

Qwen2 Instruct 72B

Alibaba

Free

Qwen2.5 Coder 32B Instruct

Alibaba

Free

Qwen2.5 Coder Instruct 7B

Alibaba

Free

Qwen2.5 Instruct 32B

Alibaba

Free

Qwen2.5 Max

Alibaba

Free

Qwen3 0.6B (Non-reasoning)

Alibaba

Free

Qwen3 0.6B (Reasoning)

Alibaba

Free

Qwen3 1.7B (Non-reasoning)

Alibaba

Free

Qwen3 1.7B (Reasoning)

Alibaba

Free

Qwen3 4B (Non-reasoning)

Alibaba

Free

Qwen3 4B (Reasoning)

Alibaba

Free

Qwen3 4B 2507 (Reasoning)

Alibaba

Free

Qwen3 4B 2507 Instruct

Alibaba

Free

Qwen3 VL 4B (Reasoning)

Alibaba

Free

Qwen3 VL 4B Instruct

Alibaba

Free

Qwen3.5 0.8B (Non-reasoning)

Alibaba

Free

Qwen3.5 0.8B (Reasoning)

Alibaba

Free

Qwen3.5 2B (Reasoning)

Alibaba

Free

QwQ 32B-Preview

Alibaba

Free

R1 1776

Perplexity

Free

Ring-1T

InclusionAI

Free

Sarvam M (Reasoning)

Sarvam

Free

Solar Open 100B (Reasoning)

Upstage

Free

Solar Pro 2 (Non-reasoning)

Upstage

Free

Solar Pro 2 (Preview) (Non-reasoning)

Upstage

Free

Solar Pro 2 (Preview) (Reasoning)

Upstage

Free

Sonar

Perplexity

Free

Sonar Reasoning

Perplexity

Free

Sonar Reasoning Pro

Perplexity

Free

Step3 VL 10B

StepFun

Free

Tiny Aya Global

Cohere

Free

Tri-21B-Think

Trillion Labs

Free

Tri-21B-think Preview

Trillion Labs

Free

Xiaomi: MiMo-V2-Omni

Xiaomi

Free

Xiaomi: MiMo-V2-Pro

Xiaomi

Free

AI API Pricing Guide

How Token-Based Pricing Works

Most LLM APIs charge per token processed, split into two categories: input tokens (the text you send — your prompt, context, and history) and outputtokens (the model's generated response). Output pricing is typically 2–4× higher than input, as generation requires more compute.

In English, 1,000 tokens correspond to roughly 750 words. A full A4 page of text contains between 600 and 900 tokens.

Real-World Monthly Cost Example

Consider a company using the GPT-4o API to process 100 emails per day, with an average prompt of 800 tokens and a 300-token response. That's 110,000 tokens/day × 30 days = 3.3 million tokens/month. At $2.50/M input tokens and $10/M output:

  • Input: 2.4M tokens × $2.50/M = $6.00/mo
  • Output: 0.9M tokens × $10/M = $9.00/mo
  • Total: $15.00/mo

The same volume with Claude Haiku (~$0.25/M input) would cost only ~$1.73/mo — a significant saving when maximum quality isn't critical.

Strategies to Reduce API Costs

1. Pick the right model for each task: simple text classification can use Gemini Flash or Claude Haiku; reserve GPT-4o or Claude Opus for tasks that truly need advanced reasoning.

2. Compress your prompts: avoid repeating unnecessary context. Well-implemented RAG systems send only the relevant passages, not the entire document.

3. Cache responses: if the same prompt is sent repeatedly (e.g., product categorization), store results and reuse them. Providers like Anthropic offer prompt caching at a discount.

4. Use open-source models via third-party APIs: Groq, Together AI, and Fireworks serve models like Llama and Qwen at $0.01–$0.20/M tokens — 10–100× cheaper than proprietary frontier models.

Frequently Asked Questions about API Costs

How much does the GPT-4o API cost per token?

The GPT-4o API costs $2.50 per million input tokens and $10.00 per million output tokens (2026 pricing). For a business sending 1 million tokens per day, the monthly cost would be approximately $75 for input alone. Output tokens are 4× more expensive, so optimizing prompt length has a significant impact on cost.

What is the cheapest AI API available?

Open source models like Qwen, Llama and Gemma can be accessed via third-party APIs (Groq, Together AI, Fireworks) for fractions of a cent per million tokens — as low as $0.01–$0.10/M tokens. Among proprietary APIs, Gemini Flash and Claude Haiku are the most affordable at $0.08–$0.25/M input tokens.

What are tokens and how do I estimate my project cost?

Tokens are text units that LLMs process — in English, 1 token ≈ 4 characters. A standard A4 page has ~600–800 tokens. To estimate cost: (input tokens + output tokens) × price/1M tokens. Example: 500-token prompt + 300-token response = 800 tokens × model price.

Should I use the API or subscribe to ChatGPT Plus/Claude Pro?

For moderate personal use, a subscription ($20/month) is usually more economical. For heavy usage or product integration, the API is more flexible and scalable. The breakeven point typically occurs when your API token consumption exceeds the equivalent value of the monthly subscription.

How have AI API prices changed over time?

Prices have dropped dramatically: GPT-4 cost $30/M tokens in 2023; equivalent models now cost $2–5/M. The trend is continuous decline as competition increases. We update this table weekly — always verify official pricing before committing your budget.

Explore the Benchmark