AI API changes: August 2026
430 changes shipped without an announcement, and 239 that came with a published date. Compiled automatically from provider catalogs and endpoints; every figure below links to the model it came from.
77
capacity cuts
context or output shrank
221
price changes
same call, different bill
24
capabilities lost
239
announced
deprecations with a date
Unannounced, biggest first
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| truncation | thedrummer | -97% max out · thedrummer/unslopnemo-12b: max_output_tokens cut 1,024,000 to 32,768 inferred | 2026-08-24 |
| truncation | openai | -97% max out · openai/gpt-4.1-nano: max_output_tokens cut 942,818 to 32,768 | 2026-08-30 |
| truncation | nvidia | -96% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 461,059 to 16,384 | 2026-08-28 |
| truncation | qwen | -94% max out · qwen/qwen3-next-80b-a3b-instruct: max_output_tokens cut 262,144 to 16,384 +1 alias | 2026-08-19 |
| cost | deepseek | 15.00x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.0036/Mtok to $0.05/Mtok | 2026-08-11 |
| cost | deepseek | 12.14x price · deepseek-v4-pro: cache_read_price_per_mtok $0.0036/Mtok to $0.04/Mtok +1 alias | 2026-08-21 |
| cost | z-ai | 11.00x price · z-ai/glm-5.2: output_price_per_mtok $0.22/Mtok to $2.42/Mtok | 2026-08-10 |
| cost | z-ai | 10.86x price · z-ai/glm-5.2: input_price_per_mtok $0.07/Mtok to $0.76/Mtok | 2026-08-10 |
| cost | z-ai | 10.77x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.01/Mtok to $0.14/Mtok | 2026-08-10 |
| truncation | novita | -91% ctx · novita/meta-llama/llama-3.3-70b-instruct: context_tokens cut 131,072 to 12,288 | 2026-08-28 |
| capability | arcee-ai | lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs · arcee-ai/trinity-large-thinking: capabilities lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs inferred | 2026-08-31 |
| cost | xai | 10.00x price · xai/grok-code-fast-1-0825: cache_read_price_per_mtok $0.02/Mtok to $0.20/Mtok +2 aliases | 2026-08-12 |
| truncation | novita | -90% max out · novita/meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 120,000 to 12,288 | 2026-08-28 |
| truncation | together_ai | -88% max out · together_ai/zai-org/GLM-5.2: max_output_tokens cut 1,048,575 to 128,000 +1 alias | 2026-08-30 |
| truncation | gemini | -88% ctx · gemini/gemini-omni-flash-preview: context_tokens cut 1,048,576 to 131,072 | 2026-08-30 |
| capability | lost frequency_penalty, logprobs, presence_penalty, repetition_penalty, stop, structured_outputs, top_k, top_logprobs · google/gemma-4-26b-a4b-it:free: capabilities lost frequency_penalty, logprobs, presence_penalty, repetition_penalty, stop, structured_outputs, top_k, top_logprobs inferred | 2026-08-20 | |
| truncation | qwen | -88% max out · qwen/qwen3-coder-30b-a3b-instruct: max_output_tokens cut 262,144 to 32,768 +2 aliases | 2026-08-13 |
| truncation | meta-llama | -87% max out · meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 128,000 to 16,384 | 2026-08-04 |
| truncation | meta | -86% max out · meta/muse-glimmer-30b: max_output_tokens cut 117,964 to 16,384 inferred | 2026-08-28 |
| cost | meta-llama | 7.10x price · meta-llama/llama-3.3-70b-instruct: input_price_per_mtok $0.10/Mtok to $0.71/Mtok | 2026-08-26 |
| cost | qwen | 6.67x price · qwen/qwen3.5-397b-a17b: cache_read_price_per_mtok $0.04/Mtok to $0.30/Mtok | 2026-08-13 |
| cost | deepseek | 6.43x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0028/Mtok to $0.02/Mtok | 2026-08-01 |
| truncation | meta-llama | -84% ctx · meta-llama/llama-guard-4-12b: context_tokens cut 1,048,576 to 163,840 | 2026-08-26 |
| cost | xai | 6.25x price · xai/grok-4-1-fast-non-reasoning-latest: input_price_per_mtok $0.20/Mtok to $1.25/Mtok +6 aliases | 2026-08-30 |
| cost | deepseek | 6.07x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.0036/Mtok to $0.02/Mtok | 2026-08-16 |
| capability | tencent | lost frequency_penalty, presence_penalty, repetition_penalty, response_format, stop, structured_outputs · tencent/hy3-preview: capabilities lost frequency_penalty, presence_penalty, repetition_penalty, response_format, stop, structured_outputs inferred | 2026-08-03 |
| truncation | deepseek | -83% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 384,000 to 65,536 inferred | 2026-08-01 |
| truncation | deepseek | -80% max out · deepseek/deepseek-v3.1-terminus: max_output_tokens cut 163,840 to 32,768 | 2026-08-20 |
| cost | xai | 5.00x price · xai/grok-3-mini-beta: output_price_per_mtok $0.50/Mtok to $2.50/Mtok +9 aliases | 2026-08-30 |
| cost | deepinfra | 5.00x price · deepinfra/Gryphe/MythoMax-L2-13b: input_price_per_mtok $0.08/Mtok to $0.40/Mtok | 2026-08-28 |
| cost | deepseek | 5.00x price · deepseek-v4-flash: cache_read_price_per_mtok $0.0028/Mtok to $0.01/Mtok +1 alias | 2026-08-21 |
| cost | xai | 5.00x price · xai/grok-code-fast-1-0825: input_price_per_mtok $0.20/Mtok to $1.00/Mtok +2 aliases | 2026-08-12 |
| cost | deepseek | 4.71x price · deepseek-v4-flash: output_price_per_mtok $0.28/Mtok to $1.32/Mtok +1 alias | 2026-08-21 |
| cost | deepseek | 4.55x price · deepseek-v4-pro: output_price_per_mtok $0.87/Mtok to $3.96/Mtok +1 alias | 2026-08-21 |
| cost | deepinfra | 4.44x price · deepinfra/Gryphe/MythoMax-L2-13b: output_price_per_mtok $0.09/Mtok to $0.40/Mtok | 2026-08-28 |
| cost | z-ai | 4.19x price · z-ai/glm-5.2: output_price_per_mtok $0.89/Mtok to $3.74/Mtok | 2026-08-03 |
| cost | z-ai | 4.19x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.05/Mtok to $0.22/Mtok | 2026-08-03 |
| cost | z-ai | 4.19x price · z-ai/glm-5.2: input_price_per_mtok $0.28/Mtok to $1.19/Mtok | 2026-08-03 |
| cost | xai | 4.17x price · xai/grok-3-mini-beta: input_price_per_mtok $0.30/Mtok to $1.25/Mtok +2 aliases | 2026-08-30 |
| truncation | qwen | -75% max out · qwen/qwen2.5-vl-72b-instruct: max_output_tokens cut 115,200 to 28,800 | 2026-08-26 |
| truncation | qwen | -75% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 262,144 to 65,536 +8 aliases | 2026-08-23 |
| truncation | ~deepseek | -75% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 1,048,576 to 262,144 +1 alias inferred | 2026-08-23 |
| truncation | liquid | -75% max out · liquid/lfm-2.5-2.6b:free: max_output_tokens cut 32,768 to 8,192 inferred | 2026-08-14 |
| cost | xai | 4.00x price · xai/grok-4-1-fast-non-reasoning-latest: cache_read_price_per_mtok $0.05/Mtok to $0.20/Mtok +6 aliases | 2026-08-30 |
| cost | gemini | 4.00x price · gemini-2.5-flash-preview-tts: output_price_per_mtok $2.50/Mtok to $10.00/Mtok +1 alias | 2026-08-28 |
| cost | deepinfra | 4.00x price · deepinfra/meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo: input_price_per_mtok $0.10/Mtok to $0.40/Mtok | 2026-08-28 |
| cost | qwen | 4.00x price · qwen/qwen3.6-27b: cache_read_price_per_mtok $0.03/Mtok to $0.12/Mtok +1 alias | 2026-08-20 |
| truncation | novita | -74% max out · novita/qwen/qwen2.5-7b-instruct: max_output_tokens cut 32,000 to 8,192 inferred | 2026-08-28 |
| cost | deepseek | 3.88x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.03/Mtok to $0.14/Mtok | 2026-08-31 |
| cost | deepseek | 3.83x price · deepseek/deepseek-v4-pro: output_price_per_mtok $0.83/Mtok to $3.20/Mtok | 2026-08-31 |
| cost | deepseek | 3.83x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.42/Mtok to $1.60/Mtok | 2026-08-31 |
| cost | nvidia | 3.53x price · nvidia/nemotron-3-super-120b-a12b: input_price_per_mtok $0.09/Mtok to $0.30/Mtok +1 alias | 2026-08-09 |
| cost | mistral | 3.33x price · mistral/mistral-small-latest: output_price_per_mtok $0.18/Mtok to $0.60/Mtok | 2026-08-21 |
| truncation | qwen | -69% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 262,144 to 81,920 +1 alias | 2026-08-24 |
| cost | deepseek | 3.14x price · deepseek-v4-flash: input_price_per_mtok $0.14/Mtok to $0.44/Mtok +1 alias | 2026-08-21 |
| cost | deepseek | 3.03x price · deepseek-v4-pro: input_price_per_mtok $0.43/Mtok to $1.32/Mtok +1 alias | 2026-08-21 |
| capability | x-ai | lost frequency_penalty, presence_penalty, stop · x-ai/grok-4.3: capabilities lost frequency_penalty, presence_penalty, stop +3 aliases | 2026-08-18 |
| capability | ~x-ai | lost frequency_penalty, presence_penalty, stop · ~x-ai/grok-latest: capabilities lost frequency_penalty, presence_penalty, stop inferred | 2026-08-18 |
| capability | anthropic | lost max_completion_tokens, response_format, structured_outputs · anthropic/claude-opus-4.1: capabilities lost max_completion_tokens, response_format, structured_outputs | 2026-08-06 |
| truncation | deepseek | -67% max out · deepseek/deepseek-v4-flash: max_output_tokens cut 393,216 to 131,072 | 2026-08-06 |
| cost | deepinfra | 3.00x price · deepinfra/Qwen/Qwen2.5-72B-Instruct: input_price_per_mtok $0.12/Mtok to $0.36/Mtok | 2026-08-28 |
| truncation | arcee-ai | -66% max out · arcee-ai/trinity-large-thinking: max_output_tokens cut 235,929 to 80,000 inferred | 2026-08-30 |
| truncation | deepseek | -66% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 384,000 to 131,072 | 2026-08-24 |
| truncation | ~deepseek | -66% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 384,000 to 131,072 inferred | 2026-08-08 |
| truncation | qwen | -65% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 235,929 to 81,920 +2 aliases | 2026-08-28 |
| cost | tencent | 2.86x price · tencent/hy3-preview: output_price_per_mtok $0.21/Mtok to $0.60/Mtok | 2026-08-15 |
| cost | tencent | 2.86x price · tencent/hy3-preview: input_price_per_mtok $0.06/Mtok to $0.18/Mtok | 2026-08-15 |
| cost | tencent | 2.86x price · tencent/hy3-preview: cache_read_price_per_mtok $0.02/Mtok to $0.06/Mtok | 2026-08-15 |
| cost | deepseek | 2.76x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.04/Mtok to $0.12/Mtok | 2026-08-19 |
| cost | xai | 2.67x price · xai/grok-3-mini-beta: cache_read_price_per_mtok $0.07/Mtok to $0.20/Mtok +2 aliases | 2026-08-30 |
| truncation | novita | -62% max out · novita/moonshotai/kimi-k2-0905: max_output_tokens cut 262,144 to 100,352 +1 alias | 2026-08-28 |
| truncation | deepseek | -61% ctx · deepseek/deepseek-r1: context_tokens cut 163,840 to 64,000 | 2026-08-13 |
| truncation | novita | -60% max out · novita/deepseek/deepseek-v3-0324: max_output_tokens cut 163,840 to 65,536 inferred | 2026-08-28 |
| truncation | deepseek | -60% max out · deepseek/deepseek-v3.2: max_output_tokens cut 163,840 to 65,536 +1 alias | 2026-08-09 |
| cost | mistral | 2.50x price · mistral/mistral-small-latest: input_price_per_mtok $0.06/Mtok to $0.15/Mtok | 2026-08-21 |
| cost | z-ai | 2.50x price · z-ai/glm-5.2: output_price_per_mtok $0.97/Mtok to $2.42/Mtok | 2026-08-17 |
| cost | z-ai | 2.47x price · z-ai/glm-5.2: input_price_per_mtok $0.31/Mtok to $0.76/Mtok | 2026-08-17 |
| truncation | deepseek | -59% max out · deepseek/deepseek-v4-pro-0813: max_output_tokens cut 943,717 to 384,000 | 2026-08-28 |
| cost | qwen | 2.45x price · qwen/qwen3-vl-235b-a22b-thinking: input_price_per_mtok $0.40/Mtok to $0.98/Mtok +1 alias | 2026-08-06 |
| cost | z-ai | 2.45x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.06/Mtok to $0.14/Mtok | 2026-08-17 |
| truncation | novita | -59% max out · novita/qwen/qwen3-4b-fp8: max_output_tokens cut 20,000 to 8,192 inferred | 2026-08-28 |
| cost | z-ai | 2.43x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-08-18 |
| cost | z-ai | 2.43x price · z-ai/glm-5.2: output_price_per_mtok $1.54/Mtok to $3.74/Mtok | 2026-08-18 |
| cost | z-ai | 2.43x price · z-ai/glm-5.2: input_price_per_mtok $0.49/Mtok to $1.19/Mtok | 2026-08-18 |
| cost | z-ai | 2.34x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-08-14 |
| cost | deepinfra | 2.33x price · deepinfra/NousResearch/Hermes-3-Llama-3.1-70B: input_price_per_mtok $0.30/Mtok to $0.70/Mtok | 2026-08-28 |
| cost | deepinfra | 2.33x price · deepinfra/NousResearch/Hermes-3-Llama-3.1-70B: output_price_per_mtok $0.30/Mtok to $0.70/Mtok | 2026-08-28 |
| cost | deepseek | 2.28x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $0.87/Mtok to $1.98/Mtok | 2026-08-16 |
| truncation | deepseek | -56% max out · deepseek/deepseek-v3.2: max_output_tokens cut 147,456 to 65,536 +2 aliases | 2026-08-30 |
| cost | nvidia | 2.25x price · nvidia/nemotron-3-super-120b-a12b: output_price_per_mtok $0.40/Mtok to $0.90/Mtok +1 alias | 2026-08-09 |
| truncation | sao10k | -55% max out · sao10k/l3-lunaris-8b: max_output_tokens cut 16,384 to 7,372 inferred | 2026-08-25 |
| cost | meta-llama | 2.22x price · meta-llama/llama-3.3-70b-instruct: output_price_per_mtok $0.32/Mtok to $0.71/Mtok | 2026-08-26 |
| cost | deepseek | 2.13x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.07/Mtok to $0.14/Mtok | 2026-08-25 |
| cost | deepseek | 2.13x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.13/Mtok to $0.28/Mtok | 2026-08-25 |
| cost | deepseek | 2.13x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-08-25 |
| cost | z-ai | 2.11x price · z-ai/glm-5.2: output_price_per_mtok $1.50/Mtok to $3.15/Mtok | 2026-08-18 |
| cost | xai | 2.08x price · xai/grok-3-mini-fast-beta: input_price_per_mtok $0.60/Mtok to $1.25/Mtok +2 aliases | 2026-08-30 |
| cost | qwen | 2.08x price · qwen/qwen3.6-27b: input_price_per_mtok $0.29/Mtok to $0.60/Mtok +1 alias | 2026-08-18 |
| truncation | z-ai | -51% max out · z-ai/glm-5.2: max_output_tokens cut 262,144 to 128,000 | 2026-08-06 |
| cost | z-ai | 2.05x price · z-ai/glm-5.2: output_price_per_mtok $1.54/Mtok to $3.15/Mtok | 2026-08-13 |
| capability | kwaipilot | lost logprobs, top_logprobs · kwaipilot/kat-coder-pro-v2.5: capabilities lost logprobs, top_logprobs +1 alias inferred | 2026-08-31 |
| capability | z-ai | lost logprobs, top_logprobs · z-ai/glm-4.7: capabilities lost logprobs, top_logprobs | 2026-08-31 |
| capability | xai | lost function_calling, tool_choice · xai/grok-4.20-multi-agent-0309: capabilities lost function_calling, tool_choice inferred | 2026-08-30 |
| truncation | qwen | -50% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 32,768 to 16,384 +2 aliases | 2026-08-28 |
| truncation | -50% ctx · google/gemma-3-27b-it: context_tokens cut 262,144 to 131,072 | 2026-08-28 | |
| capability | qwen | lost logit_bias, min_p · qwen/qwen3-235b-a22b-thinking-2507: capabilities lost logit_bias, min_p | 2026-08-24 |
| truncation | qwen | -50% ctx · qwen/qwen3-235b-a22b-thinking-2507: context_tokens cut 262,144 to 131,072 | 2026-08-24 |
| truncation | mistral | -50% max out · mistral/codestral-2508: max_output_tokens cut 256,000 to 128,000 | 2026-08-21 |
| truncation | mistral | -50% ctx · mistral/codestral-2508: context_tokens cut 256,000 to 128,000 | 2026-08-21 |
| truncation | qwen | -50% max out · qwen/qwen3.6-27b: max_output_tokens cut 131,072 to 65,536 | 2026-08-19 |
| truncation | nvidia | -50% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 262,144 to 131,072 | 2026-08-16 |
| truncation | xai | -50% max out · xai/grok-4.20-0309-reasoning: max_output_tokens cut 2,000,000 to 1,000,000 +3 aliases inferred | 2026-08-12 |
| truncation | xai | -50% ctx · xai/grok-4.20-0309-reasoning: context_tokens cut 2,000,000 to 1,000,000 +3 aliases inferred | 2026-08-12 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +5 aliases | 2026-08-31 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.96/Mtok +5 aliases | 2026-08-31 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.32/Mtok +5 aliases | 2026-08-31 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.09/Mtok to $0.18/Mtok | 2026-08-30 |
| cost | 2.00x price · ~google/gemini-flash-latest: cache_read_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-08-28 | |
| cost | 2.00x price · google/gemini-3.7-flash: output_price_per_mtok $1.88/Mtok to $3.75/Mtok | 2026-08-28 | |
| cost | 2.00x price · ~google/gemini-flash-latest: input_price_per_mtok $0.38/Mtok to $0.75/Mtok | 2026-08-28 | |
| cost | 2.00x price · ~google/gemini-flash-latest: output_price_per_mtok $1.88/Mtok to $3.75/Mtok | 2026-08-28 | |
| cost | 2.00x price · google/gemini-3.7-flash: cache_read_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-08-28 | |
| cost | 2.00x price · google/gemini-3.7-flash: input_price_per_mtok $0.38/Mtok to $0.75/Mtok | 2026-08-28 | |
| cost | deepinfra | 2.00x price · deepinfra/Qwen/Qwen3-14B: input_price_per_mtok $0.06/Mtok to $0.12/Mtok | 2026-08-28 |
| cost | gemini | 2.00x price · gemini/gemini-2.5-pro-preview-tts: output_price_per_mtok $10.00/Mtok to $20.00/Mtok | 2026-08-28 |
| cost | together_ai | 2.00x price · together_ai/Qwen/Qwen3.7-Max: output_price_per_mtok $3.75/Mtok to $7.50/Mtok | 2026-08-28 |
| cost | together_ai | 2.00x price · together_ai/Qwen/Qwen3.7-Max: input_price_per_mtok $1.25/Mtok to $2.50/Mtok | 2026-08-28 |
| cost | vertex_ai-language-models | 2.00x price · gemini-2.5-pro-preview-tts: output_price_per_mtok $10.00/Mtok to $20.00/Mtok | 2026-08-28 |
| cost | gemini | 2.00x price · gemini/gemini-3.1-flash-image-preview: input_price_per_mtok $0.25/Mtok to $0.50/Mtok +1 alias | 2026-08-21 |
| cost | gemini | 2.00x price · gemini/gemini-3.1-flash-image-preview: output_price_per_mtok $1.50/Mtok to $3.00/Mtok +1 alias | 2026-08-21 |
| cost | 2.00x price · google/gemma-4-31b-it: cache_read_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-08-21 | |
| cost | qwen | 2.00x price · qwen/qwen3.6-27b: input_price_per_mtok $0.30/Mtok to $0.60/Mtok +1 alias | 2026-08-20 |
| cost | minimax | 2.00x price · minimax/minimax-m3:batch: input_price_per_mtok $0.15/Mtok to $0.30/Mtok | 2026-08-18 |
| cost | nvidia | 2.00x price · nvidia/nemotron-3-ultra-550b-a55b:batch: input_price_per_mtok $0.30/Mtok to $0.60/Mtok | 2026-08-18 |
| cost | z-ai | 2.00x price · z-ai/glm-5.2:batch: input_price_per_mtok $0.70/Mtok to $1.40/Mtok | 2026-08-18 |
| cost | moonshotai | 2.00x price · moonshotai/kimi-k2.7-code:batch: cache_read_price_per_mtok $0.10/Mtok to $0.19/Mtok | 2026-08-18 |
| cost | z-ai | 2.00x price · z-ai/glm-5.2:batch: output_price_per_mtok $2.20/Mtok to $4.40/Mtok | 2026-08-18 |
| cost | thinkingmachines | 2.00x price · thinkingmachines/inkling:batch: cache_read_price_per_mtok $0.09/Mtok to $0.17/Mtok | 2026-08-18 |
| cost | minimax | 2.00x price · minimax/minimax-m3:batch: output_price_per_mtok $0.60/Mtok to $1.20/Mtok | 2026-08-18 |
| cost | moonshotai | 2.00x price · moonshotai/kimi-k2.7-code:batch: output_price_per_mtok $2.00/Mtok to $4.00/Mtok | 2026-08-18 |
| cost | minimax | 2.00x price · minimax/minimax-m3:batch: cache_read_price_per_mtok $0.03/Mtok to $0.06/Mtok | 2026-08-18 |
| cost | thinkingmachines | 2.00x price · thinkingmachines/inkling:batch: output_price_per_mtok $2.02/Mtok to $4.05/Mtok | 2026-08-18 |
| cost | thinkingmachines | 2.00x price · thinkingmachines/inkling:batch: input_price_per_mtok $0.50/Mtok to $1.00/Mtok | 2026-08-18 |
| cost | moonshotai | 2.00x price · moonshotai/kimi-k2.7-code:batch: input_price_per_mtok $0.47/Mtok to $0.95/Mtok | 2026-08-18 |
| cost | z-ai | 2.00x price · z-ai/glm-5.2:batch: cache_read_price_per_mtok $0.13/Mtok to $0.26/Mtok | 2026-08-18 |
| cost | nvidia | 2.00x price · nvidia/nemotron-3-ultra-550b-a55b:batch: output_price_per_mtok $1.80/Mtok to $3.60/Mtok | 2026-08-18 |
| cost | nvidia | 2.00x price · nvidia/nemotron-3-ultra-550b-a55b: cache_read_price_per_mtok $0.10/Mtok to $0.20/Mtok +1 alias | 2026-08-18 |
| cost | openai | 2.00x price · openai/gpt-5.6-luna-pro: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-terra-pro: output_price_per_mtok $6.00/Mtok to $12.00/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-terra-pro: input_price_per_mtok $1.00/Mtok to $2.00/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-terra-pro: cache_read_price_per_mtok $0.10/Mtok to $0.20/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-luna-pro: input_price_per_mtok $0.10/Mtok to $0.20/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-luna-pro: output_price_per_mtok $0.60/Mtok to $1.20/Mtok +1 alias | 2026-08-17 |
| truncation | nvidia | -49% ctx · nvidia/nemotron-3-ultra-550b-a55b: context_tokens cut 512,288 to 262,144 | 2026-08-28 |
| truncation | mistralai | -49% ctx · mistralai/mistral-small-3.2-24b-instruct: context_tokens cut 256,000 to 131,072 | 2026-08-22 |
| truncation | liquid | -49% ctx · liquid/lfm-2.5-2.6b:free: context_tokens cut 128,000 to 65,536 inferred | 2026-08-21 |
| cost | z-ai | 1.89x price · z-ai/glm-5.2: output_price_per_mtok $1.98/Mtok to $3.74/Mtok | 2026-08-14 |
| cost | z-ai | 1.89x price · z-ai/glm-5.2: input_price_per_mtok $0.63/Mtok to $1.19/Mtok | 2026-08-14 |
| cost | qwen | 1.88x price · qwen/qwen3.6-27b: input_price_per_mtok $0.32/Mtok to $0.60/Mtok +2 aliases | 2026-08-28 |
| cost | deepseek | 1.85x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-08-12 |
| cost | deepseek | 1.85x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.63/Mtok to $1.17/Mtok | 2026-08-12 |
| cost | deepseek | 1.85x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.26/Mtok to $2.34/Mtok | 2026-08-12 |
| cost | gryphe | 1.83x price · gryphe/mythomax-l2-13b: output_price_per_mtok $0.06/Mtok to $0.11/Mtok | 2026-08-04 |
| cost | qwen | 1.83x price · qwen/qwen3-vl-235b-a22b-instruct: output_price_per_mtok $1.04/Mtok to $1.90/Mtok | 2026-08-17 |
| cost | deepseek | 1.80x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok | 2026-08-20 |
| cost | qwen | 1.80x price · qwen/qwen3.6-27b: output_price_per_mtok $2.00/Mtok to $3.60/Mtok +1 alias | 2026-08-20 |
| cost | deepseek | 1.80x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.56/Mtok | 2026-08-20 |
| cost | deepseek | 1.80x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.19/Mtok | 2026-08-20 |
| truncation | nvidia | -44% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 235,929 to 131,072 | 2026-08-28 |
| cost | z-ai | 1.79x price · z-ai/glm-5.2: output_price_per_mtok $1.76/Mtok to $3.15/Mtok | 2026-08-12 |
| cost | qwen | 1.79x price · qwen/qwen3.5-35b-a3b: input_price_per_mtok $0.14/Mtok to $0.25/Mtok | 2026-08-13 |
| cost | deepseek | 1.78x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0090/Mtok to $0.02/Mtok | 2026-08-30 |
| cost | moonshotai | 1.75x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.28/Mtok to $4.00/Mtok +4 aliases | 2026-08-24 |
| cost | moonshotai | 1.75x price · moonshotai/kimi-k2.6: input_price_per_mtok $0.54/Mtok to $0.95/Mtok +4 aliases | 2026-08-24 |
| cost | moonshotai | 1.75x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.16/Mtok +4 aliases | 2026-08-24 |
| cost | deepseek | 1.75x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.08/Mtok to $0.14/Mtok +1 alias | 2026-08-24 |
| cost | ~deepseek | 1.75x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.14/Mtok to $0.25/Mtok +1 alias | 2026-08-14 |
| cost | deepseek | 1.75x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +1 alias | 2026-08-24 |
| cost | ~deepseek | 1.75x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-08-12 |
| cost | deepinfra | 1.72x price · deepinfra/Qwen/Qwen3-30B-A3B: output_price_per_mtok $0.29/Mtok to $0.50/Mtok | 2026-08-28 |
| cost | 1.71x price · google/gemma-4-26b-a4b-it: input_price_per_mtok $0.07/Mtok to $0.12/Mtok | 2026-08-10 | |
| cost | qwen | 1.70x price · qwen/qwen3.8-27b: cache_read_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-08-25 |
| cost | deepseek | 1.70x price · deepseek/deepseek-v4-pro: output_price_per_mtok $2.34/Mtok to $3.96/Mtok | 2026-08-17 |
| cost | moonshotai | 1.69x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.36/Mtok to $4.00/Mtok +2 aliases | 2026-08-19 |
| cost | moonshotai | 1.69x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.16/Mtok +2 aliases | 2026-08-19 |
| cost | moonshotai | 1.69x price · moonshotai/kimi-k2.6: input_price_per_mtok $0.56/Mtok to $0.95/Mtok +2 aliases | 2026-08-19 |
| cost | ~deepseek | 1.67x price · ~deepseek/deepseek-v4-flash-latest: input_price_per_mtok $0.03/Mtok to $0.05/Mtok | 2026-08-31 |
| cost | gemini | 1.67x price · gemini-2.5-flash-native-audio-latest: input_price_per_mtok $0.30/Mtok to $0.50/Mtok +8 aliases | 2026-08-28 |
| cost | vertex_ai-language-models | 1.67x price · gemini-live-2.5-flash-preview-native-audio-09-2025: input_price_per_mtok $0.30/Mtok to $0.50/Mtok | 2026-08-28 |
| cost | ~x-ai | 1.67x price · ~x-ai/grok-latest: cache_read_price_per_mtok $0.30/Mtok to $0.50/Mtok | 2026-08-12 |
| cost | qwen | 1.66x price · qwen/qwen3-235b-a22b-2507: input_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-08-03 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-08-16 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.44/Mtok to $4.00/Mtok +3 aliases | 2026-08-13 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: input_price_per_mtok $0.58/Mtok to $0.95/Mtok +3 aliases | 2026-08-13 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.10/Mtok to $0.16/Mtok +3 aliases | 2026-08-13 |
| cost | nvidia | 1.64x price · nvidia/nemotron-3-ultra-550b-a55b: output_price_per_mtok $2.20/Mtok to $3.60/Mtok | 2026-08-01 |
| cost | z-ai | 1.61x price · z-ai/glm-5.2: output_price_per_mtok $1.23/Mtok to $1.98/Mtok | 2026-08-14 |
| cost | z-ai | 1.61x price · z-ai/glm-5.2: input_price_per_mtok $0.39/Mtok to $0.63/Mtok | 2026-08-14 |
| cost | deepinfra | 1.60x price · deepinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning: input_price_per_mtok $0.05/Mtok to $0.08/Mtok | 2026-08-28 |
| cost | moonshotai | 1.59x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-08-15 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.18/Mtok to $0.28/Mtok | 2026-08-07 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-07 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.09/Mtok to $0.14/Mtok | 2026-08-07 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.06/Mtok to $0.09/Mtok | 2026-08-25 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.11/Mtok to $0.18/Mtok | 2026-08-25 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-08-25 |
| cost | ~deepseek | 1.58x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-11 |
| cost | ~deepseek | 1.58x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.16/Mtok to $0.25/Mtok | 2026-08-11 |
| cost | z-ai | 1.58x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.14/Mtok to $0.22/Mtok | 2026-08-17 |
| cost | ~deepseek | 1.57x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-14 |
| cost | z-ai | 1.57x price · z-ai/glm-5.2: input_price_per_mtok $0.76/Mtok to $1.19/Mtok | 2026-08-17 |
| cost | deepseek | 1.56x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.18/Mtok to $0.28/Mtok +1 alias | 2026-08-24 |
| cost | z-ai | 1.55x price · z-ai/glm-5.2: output_price_per_mtok $2.42/Mtok to $3.74/Mtok | 2026-08-17 |
| cost | qwen | 1.54x price · qwen/qwen3.5-397b-a17b: output_price_per_mtok $2.34/Mtok to $3.60/Mtok +2 aliases | 2026-08-24 |
| cost | qwen | 1.54x price · qwen/qwen3.5-122b-a10b: input_price_per_mtok $0.26/Mtok to $0.40/Mtok | 2026-08-03 |
| cost | qwen | 1.54x price · qwen/qwen3.5-122b-a10b: output_price_per_mtok $2.08/Mtok to $3.20/Mtok | 2026-08-03 |
| cost | deepseek | 1.52x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.43/Mtok to $0.66/Mtok | 2026-08-16 |
| cost | deepseek | 1.51x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-08-25 |
| cost | deepseek | 1.51x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.52/Mtok to $0.79/Mtok | 2026-08-25 |
| cost | deepseek | 1.51x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.04/Mtok to $1.58/Mtok | 2026-08-25 |
| truncation | undi95 | -33% max out · undi95/remm-slerp-l2-13b: max_output_tokens cut 6,144 to 4,096 | 2026-08-24 |
| cost | bedrock_mantle | 1.50x price · bedrock_mantle/openai.gpt-5.6-sol: output_price_per_mtok $22.00/Mtok to $33.00/Mtok | 2026-08-30 |
| cost | deepinfra | 1.50x price · deepinfra/Qwen/Qwen3-30B-A3B: input_price_per_mtok $0.08/Mtok to $0.12/Mtok | 2026-08-28 |
| cost | ~deepseek | 1.50x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.08/Mtok to $0.12/Mtok | 2026-08-25 |
| cost | qwen | 1.50x price · qwen/qwen3.6-27b: output_price_per_mtok $2.40/Mtok to $3.60/Mtok +1 alias | 2026-08-18 |
| cost | deepinfra | 1.50x price · deepinfra/google/gemma-3-12b-it: output_price_per_mtok $0.10/Mtok to $0.15/Mtok | 2026-08-28 |
| cost | moonshotai | 1.50x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.28/Mtok to $3.41/Mtok | 2026-08-16 |
| cost | deepseek | 1.48x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.43/Mtok to $0.64/Mtok | 2026-08-11 |
| cost | deepseek | 1.48x price · deepseek/deepseek-v4-pro: output_price_per_mtok $0.87/Mtok to $1.29/Mtok | 2026-08-11 |
| cost | z-ai | 1.47x price · z-ai/glm-5.1: output_price_per_mtok $2.99/Mtok to $4.40/Mtok +3 aliases | 2026-08-14 |
| cost | z-ai | 1.47x price · z-ai/glm-5.1: input_price_per_mtok $0.95/Mtok to $1.40/Mtok +3 aliases | 2026-08-14 |
| cost | z-ai | 1.47x price · z-ai/glm-5.1: cache_read_price_per_mtok $0.18/Mtok to $0.26/Mtok +3 aliases | 2026-08-14 |
| cost | minimax | 1.47x price · minimax/minimax-m2.5: input_price_per_mtok $0.15/Mtok to $0.22/Mtok | 2026-08-05 |
| truncation | ~deepseek | -32% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 384,000 to 262,144 +3 aliases inferred | 2026-08-19 |
| cost | z-ai | 1.45x price · z-ai/glm-5.2: output_price_per_mtok $1.67/Mtok to $2.42/Mtok | 2026-08-11 |
| cost | moonshotai | 1.44x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.36/Mtok to $3.41/Mtok | 2026-08-15 |
| cost | deepseek | 1.44x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-08-30 |
| cost | qwen | 1.44x price · qwen/qwen3.5-35b-a3b: output_price_per_mtok $1.25/Mtok to $1.80/Mtok +1 alias | 2026-08-27 |
| cost | z-ai | 1.43x price · z-ai/glm-5.2: input_price_per_mtok $0.53/Mtok to $0.76/Mtok | 2026-08-11 |
| cost | ~deepseek | 1.43x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0070/Mtok to $0.01/Mtok | 2026-08-30 |
| cost | deepinfra | 1.43x price · deepinfra/meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo: output_price_per_mtok $0.28/Mtok to $0.40/Mtok | 2026-08-28 |
| cost | moonshotai | 1.43x price · moonshotai/kimi-k2.5: cache_read_price_per_mtok $0.07/Mtok to $0.10/Mtok +1 alias | 2026-08-27 |
| cost | ~deepseek | 1.43x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-08-21 |
| truncation | z-ai | -30% max out · z-ai/glm-5.1: max_output_tokens cut 182,476 to 128,000 | 2026-08-30 |
| cost | z-ai | 1.42x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.10/Mtok to $0.14/Mtok | 2026-08-11 |
| cost | ~deepseek | 1.41x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-09 |
| cost | ~deepseek | 1.41x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.18/Mtok to $0.25/Mtok | 2026-08-09 |
| cost | ~deepseek | 1.40x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.10/Mtok to $0.14/Mtok | 2026-08-30 |
| cost | ~deepseek | 1.40x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-23 |
| cost | deepseek | 1.38x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.15/Mtok to $1.58/Mtok | 2026-08-26 |
| cost | deepseek | 1.38x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.05/Mtok to $0.07/Mtok | 2026-08-26 |
| cost | deepseek | 1.38x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.57/Mtok to $0.79/Mtok | 2026-08-26 |
| cost | deepseek | 1.36x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.06/Mtok to $0.08/Mtok | 2026-08-17 |
| cost | deepseek | 1.36x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.12/Mtok to $0.17/Mtok | 2026-08-17 |
| cost | deepseek | 1.36x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-08-17 |
| truncation | minimax | -26% max out · minimax/minimax-m2.7: max_output_tokens cut 176,947 to 131,072 | 2026-08-25 |
| cost | xai | 1.33x price · xai/grok-3-mini-fast-beta: cache_read_price_per_mtok $0.15/Mtok to $0.20/Mtok +2 aliases | 2026-08-30 |
| cost | deepinfra | 1.33x price · deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8: input_price_per_mtok $0.15/Mtok to $0.20/Mtok | 2026-08-28 |
| cost | deepinfra | 1.33x price · deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo: output_price_per_mtok $0.03/Mtok to $0.04/Mtok | 2026-08-28 |
| cost | deepinfra | 1.33x price · deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8: output_price_per_mtok $0.60/Mtok to $0.80/Mtok | 2026-08-28 |
| cost | ~deepseek | 1.33x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.07/Mtok to $0.10/Mtok | 2026-08-27 |
| cost | gryphe | 1.33x price · gryphe/mythomax-l2-13b: input_price_per_mtok $0.06/Mtok to $0.08/Mtok | 2026-08-04 |
| cost | moonshotai | 1.33x price · moonshotai/kimi-k2.5: input_price_per_mtok $0.45/Mtok to $0.60/Mtok +1 alias | 2026-08-27 |
| cost | moonshotai | 1.33x price · moonshotai/kimi-k2.5: output_price_per_mtok $2.25/Mtok to $3.00/Mtok +1 alias | 2026-08-27 |
| cost | xai | 1.33x price · xai/grok-code-fast-1-0825: output_price_per_mtok $1.50/Mtok to $2.00/Mtok +2 aliases | 2026-08-12 |
| cost | deepseek | 1.31x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.03/Mtok to $0.04/Mtok | 2026-08-24 |
| cost | deepseek | 1.31x price · deepseek/deepseek-v4-pro: output_price_per_mtok $0.79/Mtok to $1.04/Mtok | 2026-08-24 |
| cost | deepseek | 1.31x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.40/Mtok to $0.52/Mtok | 2026-08-24 |
| cost | deepinfra | 1.31x price · deepinfra/Sao10K/L3.1-70B-Euryale-v2.2: input_price_per_mtok $0.65/Mtok to $0.85/Mtok | 2026-08-28 |
| cost | ~deepseek | 1.31x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.02/Mtok | 2026-08-20 |
| truncation | novita | -23% max out · novita/moonshotai/kimi-k2-instruct: max_output_tokens cut 131,072 to 100,352 inferred | 2026-08-28 |
| cost | z-ai | 1.30x price · z-ai/glm-5.1: output_price_per_mtok $3.04/Mtok to $3.96/Mtok | 2026-08-25 |
| cost | z-ai | 1.30x price · z-ai/glm-5.1: input_price_per_mtok $0.97/Mtok to $1.26/Mtok | 2026-08-25 |
| cost | z-ai | 1.30x price · z-ai/glm-5.1: cache_read_price_per_mtok $0.18/Mtok to $0.23/Mtok | 2026-08-25 |
| cost | ~deepseek | 1.30x price · ~deepseek/deepseek-v4-flash-latest: input_price_per_mtok $0.06/Mtok to $0.08/Mtok | 2026-08-17 |
| cost | ~deepseek | 1.30x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.12/Mtok to $0.16/Mtok | 2026-08-17 |
| cost | ~deepseek | 1.30x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-08-17 |
| cost | z-ai | 1.30x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.09/Mtok to $0.12/Mtok | 2026-08-18 |
| cost | ~deepseek | 1.30x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.01/Mtok | 2026-08-31 |
| cost | z-ai | 1.30x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.07/Mtok to $0.09/Mtok | 2026-08-14 |
| cost | ~deepseek | 1.29x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.14/Mtok to $0.18/Mtok | 2026-08-21 |
| cost | qwen | 1.28x price · qwen/qwen3.5-397b-a17b: input_price_per_mtok $0.39/Mtok to $0.50/Mtok +2 aliases | 2026-08-24 |
| cost | novita | 1.27x price · novita/qwen/qwen3-coder-480b-a35b-instruct: input_price_per_mtok $0.30/Mtok to $0.38/Mtok | 2026-08-28 |
| truncation | thedrummer | -20% max out · thedrummer/unslopnemo-12b: max_output_tokens cut 32,768 to 26,214 inferred | 2026-08-25 |
| truncation | mistralai | -20% max out · mistralai/mistral-small-3.1-24b-instruct: max_output_tokens cut 128,000 to 102,400 | 2026-08-25 |
| truncation | nvidia | -13% max out · nvidia/nemotron-3-nano-30b-a3b: max_output_tokens cut 262,144 to 228,000 inferred | 2026-08-10 |
| truncation | azure | -12% ctx · azure/eu/gpt-5.6-luna: context_tokens cut 1,050,000 to 922,000 +11 aliases | 2026-08-21 |
| truncation | openai | -12% ctx · gpt-5.6-luna: context_tokens cut 1,050,000 to 922,000 +3 aliases | 2026-08-21 |
| truncation | openai | -10% max out · openai/gpt-3.5-turbo-0613: max_output_tokens cut 4,096 to 3,685 +1 alias | 2026-08-25 |
| truncation | gryphe | -10% max out · gryphe/mythomax-l2-13b: max_output_tokens cut 4,096 to 3,686 | 2026-08-25 |
| truncation | deepseek | -10% max out · deepseek/deepseek-r1-distill-llama-70b: max_output_tokens cut 8,192 to 7,372 | 2026-08-25 |
| truncation | microsoft | -10% max out · microsoft/phi-4: max_output_tokens cut 16,384 to 14,745 | 2026-08-25 |
| truncation | rekaai | -10% max out · rekaai/reka-edge: max_output_tokens cut 16,384 to 14,745 inferred | 2026-08-25 |
| truncation | z-ai | -10% max out · z-ai/glm-5.1: max_output_tokens cut 202,752 to 182,476 | 2026-08-25 |
| truncation | minimax | -10% max out · minimax/minimax-01: max_output_tokens cut 1,000,192 to 900,172 inferred | 2026-08-25 |
| truncation | stepfun | -10% max out · stepfun/step-3.7-flash: max_output_tokens cut 256,000 to 230,400 inferred | 2026-08-25 |
| truncation | dots-studio | -10% max out · dots-studio/dots-3-note-preview:free: max_output_tokens cut 512,000 to 460,800 inferred | 2026-08-25 |
| truncation | z-ai | -10% max out · z-ai/glm-5.2:free: max_output_tokens cut 256,000 to 230,400 inferred | 2026-08-25 |
| truncation | ibm-granite | -10% max out · ibm-granite/granite-4.0-h-micro: max_output_tokens cut 131,000 to 117,900 | 2026-08-25 |
| truncation | deepseek | -10% max out · deepseek/deepseek-chat-v3-0324: max_output_tokens cut 163,840 to 147,456 +1 alias | 2026-08-25 |
| truncation | meta-llama | -10% max out · meta-llama/llama-3.2-1b-instruct: max_output_tokens cut 60,000 to 54,000 | 2026-08-25 |
| truncation | qwen | -10% max out · qwen/qwen2.5-vl-72b-instruct: max_output_tokens cut 128,000 to 115,200 | 2026-08-25 |
| truncation | deepseek | -10% max out · deepseek/deepseek-chat-v3.1: max_output_tokens cut 161,000 to 144,900 | 2026-08-25 |
| truncation | ~moonshotai | -7% max out · ~moonshotai/kimi-latest: max_output_tokens cut 943,718 to 877,357 | 2026-08-25 |
| truncation | ~moonshotai | -7% max out · ~moonshotai/kimi-latest: max_output_tokens cut 1,048,576 to 974,842 +1 alias | 2026-08-17 |
| truncation | nvidia | -5% ctx · nvidia/nemotron-3.5-lightning: context_tokens cut 1,048,576 to 1,000,000 | 2026-08-14 |
| truncation | nvidia | -3% max out · nvidia/nemotron-3-nano-30b-a3b: max_output_tokens cut 235,929 to 228,000 inferred | 2026-08-27 |
| truncation | -1% max out · google/gemini-3.1-flash-lite-image: max_output_tokens cut 66,000 to 65,536 | 2026-08-19 | |
| capability | bedrock_converse | lost prompt_caching · global.xai.grok-4.6: capabilities lost prompt_caching +1 alias | 2026-08-30 |
| capability | qwen | lost logit_bias · qwen/qwen-2.5-7b-instruct: capabilities lost logit_bias +1 alias | 2026-08-27 |
| capability | thinkingmachines | lost structured_outputs · thinkingmachines/inkling-small: capabilities lost structured_outputs inferred | 2026-08-23 |
| capability | gemini | lost reasoning · gemini/gemini-3.1-flash-lite-image: capabilities lost reasoning | 2026-08-23 |
| capability | vertex_ai-language-models | lost reasoning · gemini-3.1-flash-lite-image: capabilities lost reasoning; gained pdf_input, video_input +1 alias | 2026-08-23 |
| capability | minimax | lost parallel_tool_calls · minimax/minimax-m2.5: capabilities lost parallel_tool_calls | 2026-08-19 |
| capability | inclusionai | lost structured_outputs · inclusionai/ling-3.0-flash: capabilities lost structured_outputs inferred | 2026-08-19 |
| capability | openrouter | lost min_p · openrouter/free: capabilities lost min_p | 2026-08-18 |
| capability | openrouter | lost max_completion_tokens · openrouter/free: capabilities lost max_completion_tokens | 2026-08-18 |
| capability | lost top_a · google/gemma-4-31b-it: capabilities lost top_a | 2026-08-15 | |
| capability | deepseek | lost max_completion_tokens · deepseek/deepseek-r1: capabilities lost max_completion_tokens | 2026-08-13 |
| capability | kwaipilot | lost frequency_penalty · kwaipilot/kat-coder-air-v2.5: capabilities lost frequency_penalty +1 alias inferred | 2026-08-13 |
| capability | openai | lost max_tokens · openai/gpt-5.2-chat: capabilities lost max_tokens | 2026-08-10 |
| capability | openai | lost max_completion_tokens · openai/gpt-5.3-chat: capabilities lost max_completion_tokens | 2026-08-08 |
| availability | kwaipilot | kwaipilot/kat-coder-air-v2.5 is no longer listed by this source — calls to it are expected to fail | 2026-08-31 |
| availability | arcee-ai | arcee-ai/virtuoso-large is no longer listed by this source — calls to it are expected to fail | 2026-08-29 |
| availability | openai | openai/gpt-5-mini:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.5-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.6-terra:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.6-luna:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | xai | xai/grok-2-1212 is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-4-turbo:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o3-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-4.1-mini:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.6-luna-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.6-terra-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-4.1:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-4.1-nano:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-3.5-turbo:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.4-nano:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o3:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.4:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.1:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | xai | xai/grok-vision-beta is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.6-sol-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-4o-mini:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.6-sol:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o1-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | xai | xai/grok-2-latest is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5-nano:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | xai | xai/grok-2-vision-1212 is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.4-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.2:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o4-mini-high:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o3-mini-high:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o4-mini:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.5:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o1:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5-codex:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-4o:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | xai | xai/grok-beta is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | xai | xai/grok-2-vision-latest is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/o3-mini:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.4-mini:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | openai | openai/gpt-5.2-pro:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | moonshotai | moonshotai/kimi-k2.7-code:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | thedrummer | thedrummer/rocinante-12b is no longer listed by this source — calls to it are expected to fail | 2026-08-28 |
| availability | z-ai | z-ai/glm-5.2:batch is no longer listed by this source — calls to it are expected to fail | 2026-08-26 |
| availability | stealth | stealth/ox-alpha is no longer listed by this source — calls to it are expected to fail | 2026-08-26 |
| availability | runwayml | runwayml/gen4_aleph is no longer listed by this source — calls to it are expected to fail | 2026-08-26 |
| availability | runwayml | runwayml/gen3a_turbo is no longer listed by this source — calls to it are expected to fail | 2026-08-26 |
| availability | google/gemma-3n-e4b-it is no longer listed by this source — calls to it are expected to fail | 2026-08-26 | |
| availability | qwen | qwen/qwen-plus-2025-07-28:thinking is no longer listed by this source — calls to it are expected to fail | 2026-08-25 |
| availability | inclusionai | inclusionai/ling-2.6-1t is no longer listed by this source — calls to it are expected to fail | 2026-08-24 |
| availability | inclusionai | inclusionai/ring-2.6-1t is no longer listed by this source — calls to it are expected to fail | 2026-08-24 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free is no longer listed by this source — calls to it are expected to fail | 2026-08-24 |
| availability | inclusionai | inclusionai/ling-2.6-flash is no longer listed by this source — calls to it are expected to fail | 2026-08-24 |
| availability | nvidia | nvidia/nemotron-nano-12b-v2-vl:free is no longer listed by this source — calls to it are expected to fail | 2026-08-24 |
| availability | nvidia | nvidia/nemotron-nano-9b-v2:free is no longer listed by this source — calls to it are expected to fail | 2026-08-24 |
| availability | openai | openai/gpt-oss-20b:free is no longer listed by this source — calls to it are expected to fail | 2026-08-22 |
| availability | deepcogito | deepcogito/cogito-v2.1-671b is no longer listed by this source — calls to it are expected to fail | 2026-08-22 |
| availability | z-ai | z-ai/glm-5.2:free is no longer listed by this source — calls to it are expected to fail | 2026-08-18 |
| availability | liquid | liquid/lfm-2.5-2.6b:free is no longer listed by this source — calls to it are expected to fail | 2026-08-18 |
| availability | gemini-3.7-flash-video-understanding-eap is no longer listed by this source — calls to it are expected to fail | 2026-08-13 | |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free is no longer listed by this source — calls to it are expected to fail | 2026-08-13 |
| availability | anthropic | claude-instant-1.0 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | claude-instant-1.2 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | claude-1.0 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | claude-instant-1.1 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | gpt-4o version 2024-05-13 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | claude-1.1 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | claude-1.2 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | claude-1.3 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | replicate | replicateopenai/gpt-oss-20b vanished; replicate/openai/gpt-oss-20b appeared with identical price and context — most likely a rename inferred | 2026-08-07 |
| availability | inclusionai | inclusionai/ling-3.0-flash:free vanished; inclusionai/ling-3.0-tiny:free appeared with identical price and context — most likely a rename inferred | 2026-08-06 |
| availability | allenai | [one catalog only — litellm still list it] allenai/olmo-3-32b-think is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-29 |
| availability | xai | [one catalog only — litellm still list it] xai/grok-2 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-28 |
| availability | xai | [one catalog only — litellm still list it] xai/grok-2-vision is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-28 |
| availability | mistralai | [one catalog only — litellm still list it] mistralai/ministral-8b is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-27 |
| availability | mancer | [one catalog only — litellm still list it] mancer/weaver is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-20 |
| availability | ai21 | [one catalog only — litellm still list it] ai21/jamba-large-1.7 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-19 |
| availability | [one catalog only — litellm still list it] imagen-4.0-fast-generate-001 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-17 | |
| availability | [one catalog only — litellm still list it] imagen-4.0-ultra-generate-001 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-17 | |
| availability | [one catalog only — litellm still list it] imagen-4.0-generate-001 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-17 | |
| availability | openai | [one catalog only — litellm still list it] openai/gpt-5.3-chat is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 |
| availability | [one catalog only — litellm still list it] gemini-robotics-er-1.5-preview is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 | |
| availability | [one catalog only — litellm still list it] gemini-3-pro-preview is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 | |
| availability | [one catalog only — litellm still list it] gemini-2.0-flash-lite is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 | |
| availability | [one catalog only — litellm still list it] gemini-2.0-flash-lite-001 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 | |
| availability | [one catalog only — litellm still list it] gemini-2.0-flash is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 | |
| availability | [one catalog only — litellm still list it] gemini-2.0-flash-001 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-10 | |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-opus-4-1-20250805 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-05 |
| availability | openai | [one catalog only — litellm still list it] gpt-4o-mini-tts-2025-03-20 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-05 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-sonnet-4-20250514 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-7-sonnet-20250219 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | [one catalog only — docs, litellm, openrouter still list it] gpt-4o-mini is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | [one catalog only — docs, litellm, openrouter still list it] o4-mini is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-sonnet-20240229 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs still list it] claude-2.1 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | [one catalog only — docs, litellm, openrouter still list it] gpt-4o is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-5-sonnet-20241022 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-opus-20240229 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | [one catalog only — docs, litellm, openrouter still list it] gpt-4.1-nano is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | [one catalog only — docs, litellm, openrouter still list it] gpt-4.1 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-5-sonnet-20240620 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-haiku-20240307 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | azure | [one catalog only — docs, litellm, openrouter still list it] gpt-4.1-mini is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-3-5-haiku-20241022 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs still list it] claude-2.0 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
| availability | anthropic | [one catalog only — docs, litellm still list it] claude-opus-4-20250514 is no longer listed by this source — calls to it are expected to fail inferred | 2026-08-01 |
Announced this period
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Pro: deprecation announced for 2026-08-27 (date already passed) +3 aliases | 2026-08-28 |
| availability | together_ai | together_ai/google/gemma-3n-E4B-it: deprecation announced for 2026-08-25 (date already passed) +1 alias | 2026-08-28 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Pro: retires 2026-08-27 (1 days ago) +3 aliases | 2026-08-28 |
| availability | together_ai | together_ai/google/gemma-3n-E4B-it: retires 2026-08-25 (3 days ago) +1 alias | 2026-08-28 |
| availability | moonshotai | moonshotai/kimi-k2.5: retirement announced for 2026-08-31 — in 4 days +1 alias | 2026-08-27 |
| availability | moonshotai | moonshotai/kimi-k2.5: retires 2026-08-31 (in 7 days) | 2026-08-24 |
| availability | azure_ai | azure_ai/deepseek-r1: deprecation announced for 2026-08-13 (date already passed) | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) inferred | 2026-08-21 |
| availability | gemini | gemini/gemini-robotics-er-1.6-preview: retires 2026-08-31 (in 10 days) | 2026-08-21 |
| availability | azure_ai | azure_ai/deepseek-r1: retires 2026-08-13 (8 days ago) | 2026-08-21 |
| availability | azure_ai | azure_ai/MAI-Image-2e: deprecation announced for 2026-08-15 (date already passed) inferred | 2026-08-21 |
| availability | gemini | gemini/gemini-robotics-er-1.6-preview: deprecation announced for 2026-08-31 — in 10 days | 2026-08-21 |
| availability | azure_ai | azure_ai/MAI-Image-2e: retires 2026-08-15 (6 days ago) inferred | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-opus-4-1: retires 2026-08-05 (16 days ago) inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) +1 alias inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-opus-4-1: retires 2026-08-05 (16 days ago) +1 alias inferred | 2026-08-21 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free: retires 2026-08-24 (in 4 days) +2 aliases inferred | 2026-08-20 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free: retirement announced for 2026-08-24 — in 4 days +2 aliases inferred | 2026-08-20 |
| availability | inclusionai | inclusionai/ling-2.6-1t: retires 2026-08-24 (in 5 days) +2 aliases inferred | 2026-08-19 |
| availability | inclusionai | inclusionai/ling-2.6-1t: retirement announced for 2026-08-24 — in 5 days +2 aliases inferred | 2026-08-19 |
| availability | deepseek | deepseek/deepseek-v3.1-terminus: retires 2026-08-17 (in 2 days) | 2026-08-15 |
| availability | deepseek | deepseek/deepseek-v3.1-terminus: retirement announced for 2026-08-17 — in 2 days | 2026-08-15 |
| availability | groq | groq/llama-3.1-8b-instant: deprecation announced for 2026-08-16 — in 3 days +1 alias inferred | 2026-08-13 |
| availability | groq | groq/meta-llama/llama-4-scout-17b-16e-instruct: deprecation announced for 2026-07-17 (date already passed) +1 alias | 2026-08-13 |
| availability | groq | groq/meta-llama/llama-4-scout-17b-16e-instruct: retires 2026-07-17 (27 days ago) +1 alias | 2026-08-13 |
| availability | groq | groq/llama-3.1-8b-instant: retires 2026-08-16 (in 3 days) +1 alias inferred | 2026-08-13 |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free: retirement announced for 2026-08-13 — in 1 days inferred | 2026-08-12 |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free: retires 2026-08-13 (in 1 days) inferred | 2026-08-12 |
| availability | mistral | mistral/devstral-2512: retires 2026-07-31 (12 days ago) +5 aliases inferred | 2026-08-12 |
| availability | anthropic | claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-sonnet-20240229-v1:0: retires 2026-07-30 (13 days ago) +8 aliases | 2026-08-12 |
| availability | gemini | gemini/imagen-4.0-fast-generate-001: retires 2026-08-17 (in 5 days) +2 aliases | 2026-08-12 |
| availability | bedrock | cohere.command-r-plus-v1:0: retires 2026-08-19 (in 7 days) +1 alias inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-haiku-20240307-v1:0: deprecation announced for 2026-09-10 — in 29 days +5 aliases | 2026-08-12 |
| availability | mistral | mistral/mistral-medium-2505: deprecation announced for 2026-08-31 — in 19 days +2 aliases inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-haiku-20240307-v1:0: retires 2026-09-10 (in 29 days) +5 aliases | 2026-08-12 |
| availability | gemini | gemini/imagen-4.0-fast-generate-001: deprecation announced for 2026-08-17 — in 5 days +2 aliases | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-sonnet-20240229-v1:0: deprecation announced for 2026-07-30 (date already passed) +8 aliases | 2026-08-12 |
| availability | mistral | mistral/mistral-medium-2505: retires 2026-08-31 (in 19 days) +2 aliases inferred | 2026-08-12 |
| availability | bedrock | cohere.command-r-plus-v1:0: deprecation announced for 2026-08-19 — in 7 days +1 alias inferred | 2026-08-12 |
| availability | anthropic | claude-opus-4-1-20250805: retires 2026-08-05 (in 4 days) +1 alias inferred | 2026-08-12 |
| availability | mistral | mistral/devstral-2512: deprecation announced for 2026-07-31 (date already passed) +5 aliases inferred | 2026-08-12 |
| availability | gemini | gemini/gemini-embedding-2-preview: retires 2026-08-10 (2 days ago) | 2026-08-12 |
| availability | gemini | gemini/gemini-embedding-2-preview: deprecation announced for 2026-08-10 (date already passed) | 2026-08-12 |
| availability | openai | gpt-5.2-chat-latest: deprecation announced for 2026-08-10 (date already passed) +1 alias | 2026-08-12 |
| availability | openai | gpt-4o-mini-search-preview-2025-03-11: deprecation announced for 2026-07-23 (date already passed) +15 aliases | 2026-08-12 |
| availability | openai | gpt-4o-audio: deprecation announced for 2026-07-20 (date already passed) +8 aliases inferred | 2026-08-05 |
| availability | openai | computer-use-preview-2025-03-11: retires 2026-07-23 (9 days ago) +17 aliases inferred | 2026-08-01 |
| availability | openai | gpt-5.2-chat-latest: retires 2026-08-10 (in 9 days) +3 aliases | 2026-08-01 |
| availability | gemini | gemini/gemini-omni-flash-preview: deprecation announced for 2026-09-30 — in 31 days | 2026-08-30 |
| availability | gemini | gemini/gemini-omni-flash-preview: retires 2026-09-30 (in 31 days) | 2026-08-30 |
| availability | azure | azure/gpt-4o-transcribe: retires 2026-10-15 (in 64 days) +1 alias inferred | 2026-08-28 |
| availability | azure | azure/gpt-4o-transcribe: deprecation announced for 2026-10-15 — in 64 days +1 alias inferred | 2026-08-28 |
| availability | anthropic | claude-sonnet-4-5-20250929: retires 2026-09-29 (in 39 days) +1 alias | 2026-08-21 |
| availability | azure | azure/eu/o3-mini-2025-01-31: retires 2026-10-01 (in 50 days) +5 aliases | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-lite: retires 2026-10-20 (in 60 days) +2 aliases | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-haiku-4-5: retires 2026-10-19 (in 59 days) +2 aliases inferred | 2026-08-21 |
| availability | azure | azure/gpt-4.1-nano-2025-04-14: retires 2026-10-14 (in 63 days) +2 aliases | 2026-08-21 |
| availability | text-completion-openai | babbage-002: retires 2026-09-28 (in 38 days) +2 aliases | 2026-08-21 |
| availability | azure | azure/o4-mini-2025-04-16: deprecation announced for 2026-10-16 — in 65 days +2 aliases | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-lite: deprecation announced for 2026-10-20 — in 60 days +2 aliases | 2026-08-21 |
| availability | azure | azure/gpt-4.1-nano: deprecation announced for 2026-10-14 — in 54 days | 2026-08-21 |
| availability | vertex_ai-video-models | vertex_ai/veo-3.1-fast-generate-001: deprecation announced for 2026-11-17 — in 88 days +1 alias inferred | 2026-08-21 |
| availability | anthropic | claude-haiku-4-5-20251001: deprecation announced for 2026-10-15 — in 55 days +1 alias | 2026-08-21 |
| availability | openai | ft-babbage-002: retires 2026-10-23 (in 83 days) +45 aliases | 2026-08-21 |
| availability | azure | azure/o4-mini-2025-04-16: retires 2026-10-16 (in 65 days) +2 aliases | 2026-08-21 |
| availability | azure | azure/gpt-image-1: retires 2026-10-23 (in 72 days) +9 aliases | 2026-08-21 |
| availability | azure | azure/eu/o1-2024-12-17: retires 2026-10-21 (in 70 days) +6 aliases | 2026-08-21 |
| availability | anthropic | claude-haiku-4-5-20251001: retires 2026-10-15 (in 55 days) +1 alias | 2026-08-21 |
| availability | openai | ft:gpt-3.5-turbo-0125: deprecation announced for 2026-10-23 — in 72 days +33 aliases | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-sonnet-4-5: deprecation announced for 2026-09-29 — in 39 days +1 alias inferred | 2026-08-21 |
| availability | azure | azure/eu/o3-mini-2025-01-31: deprecation announced for 2026-10-01 — in 50 days +4 aliases | 2026-08-21 |
| availability | anthropic | claude-sonnet-4-5-20250929: deprecation announced for 2026-09-29 — in 39 days +1 alias | 2026-08-21 |
| availability | vertex_ai-video-models | vertex_ai/veo-3.1-fast-generate-001: retires 2026-11-17 (in 88 days) +1 alias inferred | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-image: deprecation announced for 2026-10-02 — in 42 days +1 alias | 2026-08-21 |
| availability | azure | azure/eu/o1-2024-12-17: deprecation announced for 2026-10-21 — in 70 days +4 aliases | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-image: retires 2026-10-02 (in 42 days) +1 alias | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-haiku-4-5: deprecation announced for 2026-10-19 — in 59 days +2 aliases inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-haiku-4-5: deprecation announced for 2026-10-15 — in 55 days +1 alias inferred | 2026-08-21 |
| availability | text-completion-openai | babbage-002: deprecation announced for 2026-09-28 — in 38 days +2 aliases | 2026-08-21 |
Who changed the most
| Provider | Unannounced changes |
|---|---|
| deepseek | 66 |
| openai | 49 |
| z-ai | 44 |
| qwen | 30 |
| ~deepseek | 23 |
| xai | 21 |
| 20 | |
| anthropic | 20 |
| moonshotai | 20 |
| nvidia | 16 |
| deepinfra | 16 |
| novita | 8 |