Unannounced changes to AI APIs
Price moves, context-window cuts, removed capabilities and vanished models — grouped so one upstream change is one row, and ranked by how big the change is rather than when we spotted it.
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| truncation | thedrummer | -97% max out · thedrummer/unslopnemo-12b: max_output_tokens cut 1,024,000 to 32,768 inferred | 2026-08-24 |
| truncation | openai | -97% max out · openai/gpt-4.1-nano: max_output_tokens cut 942,818 to 32,768 | 2026-08-30 |
| truncation | nvidia | -96% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 461,059 to 16,384 | 2026-08-28 |
| truncation | qwen | -94% max out · qwen/qwen3-next-80b-a3b-instruct: max_output_tokens cut 262,144 to 16,384 +1 alias | 2026-08-19 |
| truncation | watsonx | -94% max out · watsonx/meta-llama/llama-4-maverick-17b: max_output_tokens cut 128,000 to 8,192 inferred | 2026-09-06 |
| truncation | openrouter | -94% max out · openrouter/anthropic/claude-sonnet-4.5: max_output_tokens cut 1,000,000 to 64,000 | 2026-09-06 |
| truncation | vertex_ai | -94% ctx · vertex_ai/xai/grok-4.1-fast-non-reasoning: context_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| truncation | vertex_ai | -94% max out · vertex_ai/xai/grok-4.1-fast-non-reasoning: max_output_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| cost | deepseek | 15.00x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.0036/Mtok to $0.05/Mtok | 2026-08-11 |
| truncation | thinkingmachines | -93% max out · thinkingmachines/inkling: max_output_tokens cut 471,859 to 32,768 | 2026-09-10 |
| truncation | qwen | -93% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 235,929 to 16,384 +5 aliases | 2026-09-09 |
| cost | deepseek | 12.14x price · deepseek-v4-pro: cache_read_price_per_mtok $0.0036/Mtok to $0.04/Mtok +1 alias | 2026-08-21 |
| cost | z-ai | 11.00x price · z-ai/glm-5.2: output_price_per_mtok $0.22/Mtok to $2.42/Mtok | 2026-08-10 |
| cost | z-ai | 10.86x price · z-ai/glm-5.2: input_price_per_mtok $0.07/Mtok to $0.76/Mtok | 2026-08-10 |
| cost | z-ai | 10.77x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.01/Mtok to $0.14/Mtok | 2026-08-10 |
| truncation | novita | -91% ctx · novita/meta-llama/llama-3.3-70b-instruct: context_tokens cut 131,072 to 12,288 | 2026-08-28 |
| cost | xai | 10.00x price · xai/grok-code-fast-1-0825: cache_read_price_per_mtok $0.02/Mtok to $0.20/Mtok +2 aliases | 2026-08-12 |
| capability | arcee-ai | lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs · arcee-ai/trinity-large-thinking: capabilities lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs inferred | 2026-08-31 |
| cost | gryphe | 10.00x price · gryphe/mythomax-l2-13b: output_price_per_mtok $0.06/Mtok to $0.60/Mtok | 2026-09-01 |
| truncation | novita | -90% max out · novita/meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 120,000 to 12,288 | 2026-08-28 |
| cost | openrouter | 9.23x price · openrouter/mistralai/mixtral-8x22b-instruct: output_price_per_mtok $0.65/Mtok to $6.00/Mtok | 2026-09-06 |
| cost | ~deepseek | 8.93x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0011/Mtok to $0.01/Mtok | 2026-09-14 |
| truncation | together_ai | -88% max out · together_ai/zai-org/GLM-5.2: max_output_tokens cut 1,048,575 to 128,000 +1 alias | 2026-08-30 |
| truncation | mistralai | -88% max out · mistralai/mistral-medium-3-5:batch: max_output_tokens cut 209,715 to 26,214 inferred | 2026-09-04 |
| truncation | fireworks_ai | -88% max out · fireworks_ai/accounts/fireworks/models/kimi-k2p5: max_output_tokens cut 262,144 to 32,768 +9 aliases inferred | 2026-07-31 |
| truncation | qwen | -88% max out · qwen/qwen3-coder-30b-a3b-instruct: max_output_tokens cut 262,144 to 32,768 +2 aliases | 2026-08-13 |
| capability | lost frequency_penalty, logprobs, presence_penalty, repetition_penalty, stop, structured_outputs, top_k, top_logprobs · google/gemma-4-26b-a4b-it:free: capabilities lost frequency_penalty, logprobs, presence_penalty, repetition_penalty, stop, structured_outputs, top_k, top_logprobs inferred | 2026-08-20 | |
| truncation | gemini | -88% ctx · gemini/gemini-omni-flash-preview: context_tokens cut 1,048,576 to 131,072 | 2026-08-30 |
| truncation | mistralai | -88% ctx · mistralai/mistral-medium-3-5:batch: context_tokens cut 262,144 to 32,768 inferred | 2026-09-04 |
| truncation | z-ai | -88% max out · z-ai/glm-4.6: max_output_tokens cut 131,072 to 16,384 +2 aliases | 2026-09-09 |
| truncation | vertex_ai-language-models | -88% ctx · gemini-live-2.5-flash-native-audio: context_tokens cut 1,048,576 to 131,072 inferred | 2026-09-10 |
| truncation | meta-llama | -87% max out · meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 128,000 to 16,384 | 2026-08-04 |
| truncation | ~z-ai | -86% max out · ~z-ai/glm-latest: max_output_tokens cut 943,718 to 128,000 +1 alias inferred | 2026-09-11 |
| truncation | qwen | -86% max out · qwen/qwen3-30b-a3b-instruct-2507: max_output_tokens cut 235,929 to 32,000 +2 aliases | 2026-09-13 |
| truncation | deepseek | -86% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 943,718 to 131,072 | 2026-09-06 |
| truncation | ~z-ai | -86% max out · ~z-ai/glm-flash-latest: max_output_tokens cut 943,718 to 131,072 +3 aliases inferred | 2026-09-09 |
| truncation | z-ai | -86% max out · z-ai/glm-5.3: max_output_tokens cut 943,718 to 131,072 | 2026-09-13 |
| truncation | ~deepseek | -86% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 943,718 to 131,072 +1 alias inferred | 2026-09-13 |
| truncation | qwen | -86% max out · qwen/qwen3-next-80b-a3b-thinking: max_output_tokens cut 235,929 to 32,768 | 2026-09-15 |
| truncation | meta | -86% max out · meta/muse-glimmer-30b: max_output_tokens cut 117,964 to 16,384 inferred | 2026-08-28 |
| cost | meta-llama | 7.10x price · meta-llama/llama-3.3-70b-instruct: input_price_per_mtok $0.10/Mtok to $0.71/Mtok | 2026-08-26 |
| truncation | meta-llama | -86% max out · meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 115,200 to 16,384 +1 alias | 2026-09-14 |
| cost | qwen | 6.67x price · qwen/qwen3.5-397b-a17b: cache_read_price_per_mtok $0.04/Mtok to $0.30/Mtok | 2026-08-13 |
| cost | gryphe | 6.67x price · gryphe/mythomax-l2-13b: input_price_per_mtok $0.06/Mtok to $0.40/Mtok | 2026-09-01 |
| cost | deepseek | 6.43x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0028/Mtok to $0.02/Mtok | 2026-08-01 |
| truncation | meta-llama | -84% ctx · meta-llama/llama-guard-4-12b: context_tokens cut 1,048,576 to 163,840 | 2026-08-26 |
| cost | xai | 6.25x price · xai/grok-4-1-fast-non-reasoning-latest: input_price_per_mtok $0.20/Mtok to $1.25/Mtok +6 aliases | 2026-08-30 |
| cost | deepseek | 6.07x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.0036/Mtok to $0.02/Mtok | 2026-08-16 |
| capability | tencent | lost frequency_penalty, presence_penalty, repetition_penalty, response_format, stop, structured_outputs · tencent/hy3-preview: capabilities lost frequency_penalty, presence_penalty, repetition_penalty, response_format, stop, structured_outputs inferred | 2026-08-03 |
| truncation | deepseek | -83% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 384,000 to 65,536 inferred | 2026-08-01 |
| truncation | nvidia | -82% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 182,520 to 32,768 +1 alias | 2026-09-16 |
| cost | openrouter | 5.56x price · openrouter/qwen/qwen-2.5-coder-32b-instruct: output_price_per_mtok $0.18/Mtok to $1.00/Mtok | 2026-09-06 |
| cost | xai | 5.00x price · xai/grok-code-fast-1-0825: input_price_per_mtok $0.20/Mtok to $1.00/Mtok +2 aliases | 2026-08-12 |
| truncation | deepseek | -80% max out · deepseek/deepseek-v3.1-terminus: max_output_tokens cut 163,840 to 32,768 | 2026-08-20 |
| cost | deepseek | 5.00x price · deepseek-v4-flash: cache_read_price_per_mtok $0.0028/Mtok to $0.01/Mtok +1 alias | 2026-08-21 |
| cost | deepinfra | 5.00x price · deepinfra/Gryphe/MythoMax-L2-13b: input_price_per_mtok $0.08/Mtok to $0.40/Mtok | 2026-08-28 |
| cost | xai | 5.00x price · xai/grok-3-mini-beta: output_price_per_mtok $0.50/Mtok to $2.50/Mtok +9 aliases | 2026-08-30 |
| truncation | openrouter | -80% ctx · openrouter/anthropic/claude-sonnet-4.5: context_tokens cut 1,000,000 to 200,000 | 2026-09-06 |
| cost | ~deepseek | 5.00x price · ~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.0030/Mtok to $0.01/Mtok | 2026-09-15 |
| cost | ~deepseek | 4.77x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.02/Mtok to $0.09/Mtok +1 alias | 2026-09-16 |
| cost | deepseek | 4.71x price · deepseek-v4-flash: output_price_per_mtok $0.28/Mtok to $1.32/Mtok +1 alias | 2026-08-21 |
| cost | deepseek | 4.55x price · deepseek-v4-pro: output_price_per_mtok $0.87/Mtok to $3.96/Mtok +1 alias | 2026-08-21 |
| cost | deepinfra | 4.44x price · deepinfra/Gryphe/MythoMax-L2-13b: output_price_per_mtok $0.09/Mtok to $0.40/Mtok | 2026-08-28 |
| truncation | deepseek | -77% max out · deepseek/deepseek-chat-v3.1: max_output_tokens cut 144,900 to 32,768 +1 alias | 2026-09-08 |
| cost | deepseek | 4.23x price · deepseek/deepseek-chat-v3.1: cache_read_price_per_mtok $0.13/Mtok to $0.55/Mtok +1 alias | 2026-09-03 |
| cost | z-ai | 4.19x price · z-ai/glm-5.2: output_price_per_mtok $0.89/Mtok to $3.74/Mtok | 2026-08-03 |
| cost | z-ai | 4.19x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.05/Mtok to $0.22/Mtok | 2026-08-03 |
| cost | z-ai | 4.19x price · z-ai/glm-5.2: input_price_per_mtok $0.28/Mtok to $1.19/Mtok | 2026-08-03 |
| cost | xai | 4.17x price · xai/grok-3-mini-beta: input_price_per_mtok $0.30/Mtok to $1.25/Mtok +2 aliases | 2026-08-30 |
| cost | fireworks_ai | 4.14x price · fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: cache_read_price_per_mtok $0.14/Mtok to $0.60/Mtok +1 alias | 2026-09-15 |
| truncation | nebius | -76% ctx · nebius/Qwen/Qwen2.5-VL-72B-Instruct: context_tokens cut 131,072 to 32,000 | 2026-09-06 |
| truncation | nebius | -76% max out · nebius/Qwen/Qwen2.5-VL-72B-Instruct: max_output_tokens cut 131,072 to 32,000 | 2026-09-06 |
| truncation | ~z-ai | -75% max out · ~z-ai/glm-latest: max_output_tokens cut 943,718 to 235,929 +3 aliases inferred | 2026-09-15 |
| truncation | fireworks_ai | -75% max out · fireworks_ai/accounts/fireworks/models/gpt-oss-120b: max_output_tokens cut 131,072 to 32,768 +1 alias | 2026-06-18 |
| cost | cohere_chat | 4.00x price · command-r7b-12-2024: output_price_per_mtok $0.04/Mtok to $0.15/Mtok | 2026-06-18 |
| truncation | liquid | -75% max out · liquid/lfm-2.5-2.6b:free: max_output_tokens cut 32,768 to 8,192 inferred | 2026-08-14 |
| cost | qwen | 4.00x price · qwen/qwen3.6-27b: cache_read_price_per_mtok $0.03/Mtok to $0.12/Mtok +1 alias | 2026-08-20 |
| truncation | ~deepseek | -75% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 1,048,576 to 262,144 +1 alias inferred | 2026-08-23 |
| truncation | qwen | -75% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 262,144 to 65,536 +8 aliases | 2026-08-23 |
| truncation | qwen | -75% max out · qwen/qwen2.5-vl-72b-instruct: max_output_tokens cut 115,200 to 28,800 | 2026-08-26 |
| cost | gemini | 4.00x price · gemini-2.5-flash-preview-tts: output_price_per_mtok $2.50/Mtok to $10.00/Mtok +1 alias | 2026-08-28 |
| cost | deepinfra | 4.00x price · deepinfra/meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo: input_price_per_mtok $0.10/Mtok to $0.40/Mtok | 2026-08-28 |
| cost | xai | 4.00x price · xai/grok-4-1-fast-non-reasoning-latest: cache_read_price_per_mtok $0.05/Mtok to $0.20/Mtok +6 aliases | 2026-08-30 |
| truncation | openai | -75% ctx · gpt-realtime-mini: context_tokens cut 128,000 to 32,000 | 2026-09-03 |
| capability | vertex_ai-language-models | lost pdf_input, prompt_caching, response_schema, url_context · gemini-live-2.5-flash-native-audio: capabilities lost pdf_input, prompt_caching, response_schema, url_context inferred | 2026-09-10 |
| truncation | novita | -74% max out · novita/qwen/qwen2.5-7b-instruct: max_output_tokens cut 32,000 to 8,192 inferred | 2026-08-28 |
| cost | qwen | 3.91x price · qwen/qwen3.5-35b-a3b: input_price_per_mtok $0.08/Mtok to $0.31/Mtok | 2026-09-05 |
| cost | deepseek | 3.88x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.03/Mtok to $0.14/Mtok | 2026-08-31 |
| truncation | azure | -74% ctx · azure/gpt-5.4-mini-2026-03-17: context_tokens cut 1,050,000 to 272,000 +3 aliases | 2026-07-30 |
| truncation | openai | -74% ctx · gpt-5.4-mini-2026-03-17: context_tokens cut 1,050,000 to 272,000 +3 aliases | 2026-07-30 |
| cost | deepseek | 3.83x price · deepseek/deepseek-v4-pro: output_price_per_mtok $0.83/Mtok to $3.20/Mtok | 2026-08-31 |
| cost | deepseek | 3.83x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.42/Mtok to $1.60/Mtok | 2026-08-31 |
| cost | openrouter | 3.83x price · openrouter/qwen/qwen3-235b-a22b-thinking-2507: output_price_per_mtok $0.60/Mtok to $2.30/Mtok | 2026-09-06 |
| truncation | nvidia | -74% ctx · nvidia/nemotron-3-super-120b-a12b: context_tokens cut 1,000,000 to 262,144 +1 alias | 2026-09-09 |
| cost | qwen | 3.79x price · qwen/qwen3-14b: output_price_per_mtok $0.24/Mtok to $0.91/Mtok +1 alias | 2026-09-08 |
| cost | openrouter | 3.79x price · openrouter/qwen/qwen3-14b: output_price_per_mtok $0.24/Mtok to $0.91/Mtok | 2026-09-13 |
| cost | mistral | 3.75x price · mistral/mistral-medium-latest: input_price_per_mtok $0.40/Mtok to $1.50/Mtok | 2026-06-26 |
| cost | mistral | 3.75x price · mistral/mistral-medium-latest: output_price_per_mtok $2.00/Mtok to $7.50/Mtok | 2026-06-26 |
| cost | z-ai | 3.72x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.05/Mtok to $0.19/Mtok +1 alias | 2026-09-11 |
| cost | openrouter | 3.67x price · openrouter/qwen/qwen-2.5-coder-32b-instruct: input_price_per_mtok $0.18/Mtok to $0.66/Mtok | 2026-09-06 |
| truncation | qwen | -72% max out · qwen/qwen3.5-35b-a3b: max_output_tokens cut 235,929 to 65,536 +3 aliases | 2026-09-06 |
| cost | openrouter | 3.57x price · openrouter/deepseek/deepseek-chat-v3-0324: output_price_per_mtok $0.28/Mtok to $1.00/Mtok | 2026-09-06 |
| cost | nvidia | 3.53x price · nvidia/nemotron-3-super-120b-a12b: input_price_per_mtok $0.09/Mtok to $0.30/Mtok +1 alias | 2026-08-09 |
| cost | openrouter | 3.51x price · openrouter/mistralai/mistral-small-3.1-24b-instruct: input_price_per_mtok $0.10/Mtok to $0.35/Mtok | 2026-09-06 |
| cost | openrouter | 3.50x price · openrouter/qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.10/Mtok to $0.35/Mtok | 2026-09-06 |
| cost | z-ai | 3.45x price · z-ai/glm-5.2: output_price_per_mtok $0.88/Mtok to $3.04/Mtok +1 alias | 2026-09-11 |
| cost | z-ai | 3.45x price · z-ai/glm-5.2: input_price_per_mtok $0.28/Mtok to $0.97/Mtok +1 alias | 2026-09-11 |
| cost | mistral | 3.33x price · mistral/mistral-small-latest: output_price_per_mtok $0.18/Mtok to $0.60/Mtok | 2026-08-21 |
| cost | openrouter | 3.33x price · openrouter/mistralai/devstral-2512: output_price_per_mtok $0.60/Mtok to $2.00/Mtok | 2026-09-06 |
| truncation | qwen | -69% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 262,144 to 81,920 +1 alias | 2026-08-24 |
| cost | qwen | 3.20x price · qwen/qwen2.5-vl-72b-instruct: input_price_per_mtok $0.25/Mtok to $0.80/Mtok +1 alias | 2026-09-03 |
| cost | bedrock | 3.20x price · eu.anthropic.claude-3-5-haiku-20241022-v1:0: input_price_per_mtok $0.25/Mtok to $0.80/Mtok | 2026-09-10 |
| cost | bedrock | 3.20x price · eu.anthropic.claude-3-5-haiku-20241022-v1:0: output_price_per_mtok $1.25/Mtok to $4.00/Mtok | 2026-09-10 |
| cost | bedrock | 3.20x price · eu.anthropic.claude-3-5-haiku-20241022-v1:0: cache_read_price_per_mtok $0.03/Mtok to $0.08/Mtok | 2026-09-10 |
| cost | openrouter | 3.18x price · openrouter/deepseek/deepseek-chat: output_price_per_mtok $0.28/Mtok to $0.89/Mtok | 2026-09-06 |
| cost | deepseek | 3.14x price · deepseek-v4-flash: input_price_per_mtok $0.14/Mtok to $0.44/Mtok +1 alias | 2026-08-21 |
| truncation | openrouter | -68% max out · openrouter/anthropic/claude-haiku-4.5: max_output_tokens cut 200,000 to 64,000 | 2026-09-06 |
| cost | openrouter | 3.08x price · openrouter/mistralai/mixtral-8x22b-instruct: input_price_per_mtok $0.65/Mtok to $2.00/Mtok | 2026-09-06 |
| cost | deepseek | 3.03x price · deepseek-v4-pro: input_price_per_mtok $0.43/Mtok to $1.32/Mtok +1 alias | 2026-08-21 |
| capability | gemini | lost code_execution, file_search, service_tier · gemini/gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search, service_tier +1 alias | 2026-06-27 |
| capability | vertex_ai-language-models | lost code_execution, file_search, service_tier · gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search, service_tier +2 aliases | 2026-06-27 |
| truncation | deepseek | -67% max out · deepseek/deepseek-v4-flash: max_output_tokens cut 393,216 to 131,072 | 2026-08-06 |
| capability | anthropic | lost max_completion_tokens, response_format, structured_outputs · anthropic/claude-opus-4.1: capabilities lost max_completion_tokens, response_format, structured_outputs | 2026-08-06 |
| capability | x-ai | lost frequency_penalty, presence_penalty, stop · x-ai/grok-4.3: capabilities lost frequency_penalty, presence_penalty, stop +3 aliases | 2026-08-18 |
| capability | ~x-ai | lost frequency_penalty, presence_penalty, stop · ~x-ai/grok-latest: capabilities lost frequency_penalty, presence_penalty, stop inferred | 2026-08-18 |
| cost | deepinfra | 3.00x price · deepinfra/Qwen/Qwen2.5-72B-Instruct: input_price_per_mtok $0.12/Mtok to $0.36/Mtok | 2026-08-28 |
| truncation | ~deepseek | -67% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 393,216 to 131,072 +1 alias inferred | 2026-09-04 |
| capability | nvidia | lost logprobs, structured_outputs, top_logprobs · nvidia/nemotron-3-super-120b-a12b: capabilities lost logprobs, structured_outputs, top_logprobs | 2026-09-09 |
| cost | upstage | 3.00x price · upstage/solar-pro4: input_price_per_mtok $0.03/Mtok to $0.09/Mtok | 2026-09-11 |
| cost | upstage | 3.00x price · upstage/solar-pro4: output_price_per_mtok $0.12/Mtok to $0.36/Mtok | 2026-09-11 |
| capability | qwen | lost logit_bias, min_p, repetition_penalty · qwen/qwen3.8-flash: capabilities lost logit_bias, min_p, repetition_penalty | 2026-09-15 |
| cost | upstage | 3.00x price · upstage/solar-pro4: cache_read_price_per_mtok $0.0060/Mtok to $0.02/Mtok | 2026-09-11 |
| truncation | arcee-ai | -66% max out · arcee-ai/trinity-large-thinking: max_output_tokens cut 235,929 to 80,000 inferred | 2026-08-30 |
| truncation | ~deepseek | -66% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 384,000 to 131,072 inferred | 2026-08-08 |
| truncation | deepseek | -66% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 384,000 to 131,072 | 2026-08-24 |
| truncation | qwen | -65% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 235,929 to 81,920 +2 aliases | 2026-08-28 |
| cost | tencent | 2.86x price · tencent/hy3-preview: output_price_per_mtok $0.21/Mtok to $0.60/Mtok | 2026-08-15 |
| cost | tencent | 2.86x price · tencent/hy3-preview: input_price_per_mtok $0.06/Mtok to $0.18/Mtok | 2026-08-15 |
| cost | tencent | 2.86x price · tencent/hy3-preview: cache_read_price_per_mtok $0.02/Mtok to $0.06/Mtok | 2026-08-15 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.10/Mtok to $0.28/Mtok | 2026-09-07 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0100/Mtok to $0.03/Mtok | 2026-09-07 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.05/Mtok to $0.14/Mtok | 2026-09-07 |
| cost | jina_ai | 2.78x price · jina-reranker-v2-base-multilingual: input_price_per_mtok $0.02/Mtok to $0.05/Mtok | 2026-09-11 |
| cost | deepseek | 2.76x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.04/Mtok to $0.12/Mtok | 2026-08-19 |
| cost | xai | 2.67x price · xai/grok-3-mini-beta: cache_read_price_per_mtok $0.07/Mtok to $0.20/Mtok +2 aliases | 2026-08-30 |
| cost | openrouter | 2.67x price · openrouter/mistralai/devstral-2512: input_price_per_mtok $0.15/Mtok to $0.40/Mtok | 2026-09-06 |
| cost | deepseek | 2.63x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.06/Mtok | 2026-09-08 |
| truncation | novita | -62% max out · novita/moonshotai/kimi-k2-0905: max_output_tokens cut 262,144 to 100,352 +1 alias | 2026-08-28 |
| truncation | deepseek | -61% ctx · deepseek/deepseek-r1: context_tokens cut 163,840 to 64,000 | 2026-08-13 |
| cost | openrouter | 2.51x price · openrouter/qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.35/Mtok to $0.88/Mtok | 2026-09-13 |
| cost | openrouter | 2.51x price · openrouter/qwen/qwen3-235b-a22b-2507: input_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-09-13 |
| truncation | deepseek | -60% max out · deepseek/deepseek-v3.2: max_output_tokens cut 163,840 to 65,536 +1 alias | 2026-08-09 |
| cost | z-ai | 2.50x price · z-ai/glm-5.2: output_price_per_mtok $0.97/Mtok to $2.42/Mtok | 2026-08-17 |
| cost | mistral | 2.50x price · mistral/mistral-small-latest: input_price_per_mtok $0.06/Mtok to $0.15/Mtok | 2026-08-21 |
| truncation | novita | -60% max out · novita/deepseek/deepseek-v3-0324: max_output_tokens cut 163,840 to 65,536 inferred | 2026-08-28 |
| truncation | openrouter | -60% max out · openrouter/deepseek/deepseek-v3.2: max_output_tokens cut 163,840 to 65,536 | 2026-09-13 |
| cost | z-ai | 2.47x price · z-ai/glm-5.2: input_price_per_mtok $0.31/Mtok to $0.76/Mtok | 2026-08-17 |
| truncation | deepseek | -59% max out · deepseek/deepseek-v4-pro-0813: max_output_tokens cut 943,718 to 384,000 | 2026-09-11 |
| truncation | deepseek | -59% max out · deepseek/deepseek-v4-pro-0813: max_output_tokens cut 943,717 to 384,000 | 2026-08-28 |
| cost | qwen | 2.45x price · qwen/qwen3-vl-235b-a22b-thinking: input_price_per_mtok $0.40/Mtok to $0.98/Mtok +1 alias | 2026-08-06 |
| cost | z-ai | 2.45x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.06/Mtok to $0.14/Mtok | 2026-08-17 |
| cost | qwen | 2.44x price · qwen/qwen3-235b-a22b-2507: input_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-09-09 |
| truncation | novita | -59% max out · novita/qwen/qwen3-4b-fp8: max_output_tokens cut 20,000 to 8,192 inferred | 2026-08-28 |
| cost | z-ai | 2.43x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-08-18 |
| cost | z-ai | 2.43x price · z-ai/glm-5.2: output_price_per_mtok $1.54/Mtok to $3.74/Mtok | 2026-08-18 |
| cost | z-ai | 2.43x price · z-ai/glm-5.2: input_price_per_mtok $0.49/Mtok to $1.19/Mtok | 2026-08-18 |
| truncation | ~deepseek | -58% max out · ~deepseek/deepseek-flash-latest: max_output_tokens cut 943,718 to 393,216 +2 aliases inferred | 2026-09-15 |
| cost | fireworks_ai | 2.36x price · fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: output_price_per_mtok $0.28/Mtok to $0.66/Mtok +1 alias | 2026-09-01 |
| truncation | moonshotai | -57% max out · moonshotai/kimi-k2-thinking: max_output_tokens cut 235,929 to 100,352 +2 aliases | 2026-09-13 |
| cost | ~z-ai | 2.34x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.10/Mtok to $0.23/Mtok | 2026-09-06 |
| cost | z-ai | 2.34x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-08-14 |
| cost | deepinfra | 2.33x price · deepinfra/NousResearch/Hermes-3-Llama-3.1-70B: input_price_per_mtok $0.30/Mtok to $0.70/Mtok | 2026-08-28 |
| cost | deepinfra | 2.33x price · deepinfra/NousResearch/Hermes-3-Llama-3.1-70B: output_price_per_mtok $0.30/Mtok to $0.70/Mtok | 2026-08-28 |
| cost | openrouter | 2.29x price · openrouter/deepseek/deepseek-chat: input_price_per_mtok $0.14/Mtok to $0.32/Mtok | 2026-09-06 |
| cost | deepseek | 2.28x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $0.87/Mtok to $1.98/Mtok | 2026-08-16 |
| cost | nvidia | 2.25x price · nvidia/nemotron-3-super-120b-a12b: output_price_per_mtok $0.40/Mtok to $0.90/Mtok +1 alias | 2026-08-09 |
| truncation | deepseek | -56% max out · deepseek/deepseek-v3.2: max_output_tokens cut 147,456 to 65,536 +2 aliases | 2026-08-30 |
| truncation | sao10k | -55% max out · sao10k/l3-lunaris-8b: max_output_tokens cut 16,384 to 7,372 inferred | 2026-08-25 |
| cost | meta-llama | 2.22x price · meta-llama/llama-3.3-70b-instruct: output_price_per_mtok $0.32/Mtok to $0.71/Mtok | 2026-08-26 |
| cost | deepseek | 2.20x price · deepseek/deepseek-chat-v3.1: input_price_per_mtok $0.25/Mtok to $0.55/Mtok +1 alias | 2026-09-03 |
| cost | deepseek | 2.16x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.06/Mtok to $0.14/Mtok | 2026-09-12 |
| cost | 2.14x price · google/gemma-4-26b-a4b-it: input_price_per_mtok $0.04/Mtok to $0.09/Mtok | 2026-09-13 | |
| cost | deepseek | 2.14x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.50/Mtok to $3.20/Mtok | 2026-09-12 |
| cost | deepseek | 2.14x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.75/Mtok to $1.60/Mtok | 2026-09-12 |
| cost | deepseek | 2.13x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.07/Mtok to $0.14/Mtok | 2026-08-25 |
| cost | deepseek | 2.13x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.13/Mtok to $0.28/Mtok | 2026-08-25 |
| cost | deepseek | 2.13x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-08-25 |
| truncation | openai | -53% max out · gpt-5-pro-2025-10-06: max_output_tokens cut 272,000 to 128,000 +1 alias | 2026-06-26 |
| cost | z-ai | 2.11x price · z-ai/glm-5.2: output_price_per_mtok $1.50/Mtok to $3.15/Mtok | 2026-08-18 |
| cost | openrouter | 2.09x price · openrouter/qwen/qwen3-235b-a22b-thinking-2507: input_price_per_mtok $0.11/Mtok to $0.23/Mtok | 2026-09-06 |
| cost | xai | 2.08x price · xai/grok-3-mini-fast-beta: input_price_per_mtok $0.60/Mtok to $1.25/Mtok +2 aliases | 2026-08-30 |
| cost | qwen | 2.08x price · qwen/qwen3.6-27b: input_price_per_mtok $0.29/Mtok to $0.60/Mtok +1 alias | 2026-08-18 |
| cost | z-ai | 2.05x price · z-ai/glm-5.2: output_price_per_mtok $2.15/Mtok to $4.40/Mtok | 2026-09-15 |
| cost | z-ai | 2.05x price · z-ai/glm-5.2: input_price_per_mtok $0.68/Mtok to $1.40/Mtok | 2026-09-15 |
| truncation | z-ai | -51% max out · z-ai/glm-5.2: max_output_tokens cut 262,144 to 128,000 | 2026-08-06 |
| cost | z-ai | 2.05x price · z-ai/glm-5.2: output_price_per_mtok $1.54/Mtok to $3.15/Mtok | 2026-08-13 |
| cost | bedrock_converse | 2.05x price · qwen.qwen3-coder-480b-a35b-v1:0: input_price_per_mtok $0.22/Mtok to $0.45/Mtok | 2026-09-06 |
| cost | sambanova | 2.00x price · sambanova/MiniMax-M2.7: output_price_per_mtok $1.20/Mtok to $2.40/Mtok | 2026-06-26 |
| cost | sambanova | 2.00x price · sambanova/MiniMax-M2.7: input_price_per_mtok $0.30/Mtok to $0.60/Mtok | 2026-06-26 |
| capability | vertex_ai-language-models | lost code_execution, file_search · vertex_ai/gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search | 2026-06-27 |
| capability | openrouter | lost code_execution, file_search · openrouter/google/gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search +1 alias | 2026-06-27 |
| truncation | xai | -50% max out · xai/grok-4.20-0309-reasoning: max_output_tokens cut 2,000,000 to 1,000,000 +3 aliases inferred | 2026-08-12 |
| truncation | xai | -50% ctx · xai/grok-4.20-0309-reasoning: context_tokens cut 2,000,000 to 1,000,000 +3 aliases inferred | 2026-08-12 |
| truncation | nvidia | -50% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 262,144 to 131,072 | 2026-08-16 |
| cost | openai | 2.00x price · openai/gpt-5.6-luna-pro: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-terra-pro: output_price_per_mtok $6.00/Mtok to $12.00/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-terra-pro: input_price_per_mtok $1.00/Mtok to $2.00/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-terra-pro: cache_read_price_per_mtok $0.10/Mtok to $0.20/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-luna-pro: input_price_per_mtok $0.10/Mtok to $0.20/Mtok +1 alias | 2026-08-17 |
| cost | openai | 2.00x price · openai/gpt-5.6-luna-pro: output_price_per_mtok $0.60/Mtok to $1.20/Mtok +1 alias | 2026-08-17 |
| cost | minimax | 2.00x price · minimax/minimax-m3:batch: input_price_per_mtok $0.15/Mtok to $0.30/Mtok | 2026-08-18 |
| cost | nvidia | 2.00x price · nvidia/nemotron-3-ultra-550b-a55b:batch: input_price_per_mtok $0.30/Mtok to $0.60/Mtok | 2026-08-18 |
| cost | z-ai | 2.00x price · z-ai/glm-5.2:batch: input_price_per_mtok $0.70/Mtok to $1.40/Mtok | 2026-08-18 |
| cost | moonshotai | 2.00x price · moonshotai/kimi-k2.7-code:batch: cache_read_price_per_mtok $0.10/Mtok to $0.19/Mtok | 2026-08-18 |
| cost | z-ai | 2.00x price · z-ai/glm-5.2:batch: output_price_per_mtok $2.20/Mtok to $4.40/Mtok | 2026-08-18 |
| cost | thinkingmachines | 2.00x price · thinkingmachines/inkling:batch: cache_read_price_per_mtok $0.09/Mtok to $0.17/Mtok | 2026-08-18 |
| cost | minimax | 2.00x price · minimax/minimax-m3:batch: output_price_per_mtok $0.60/Mtok to $1.20/Mtok | 2026-08-18 |
| cost | moonshotai | 2.00x price · moonshotai/kimi-k2.7-code:batch: output_price_per_mtok $2.00/Mtok to $4.00/Mtok | 2026-08-18 |
| cost | minimax | 2.00x price · minimax/minimax-m3:batch: cache_read_price_per_mtok $0.03/Mtok to $0.06/Mtok | 2026-08-18 |
| cost | thinkingmachines | 2.00x price · thinkingmachines/inkling:batch: output_price_per_mtok $2.02/Mtok to $4.05/Mtok | 2026-08-18 |
| cost | thinkingmachines | 2.00x price · thinkingmachines/inkling:batch: input_price_per_mtok $0.50/Mtok to $1.00/Mtok | 2026-08-18 |
| cost | moonshotai | 2.00x price · moonshotai/kimi-k2.7-code:batch: input_price_per_mtok $0.47/Mtok to $0.95/Mtok | 2026-08-18 |
| cost | z-ai | 2.00x price · z-ai/glm-5.2:batch: cache_read_price_per_mtok $0.13/Mtok to $0.26/Mtok | 2026-08-18 |
| cost | nvidia | 2.00x price · nvidia/nemotron-3-ultra-550b-a55b:batch: output_price_per_mtok $1.80/Mtok to $3.60/Mtok | 2026-08-18 |
| cost | nvidia | 2.00x price · nvidia/nemotron-3-ultra-550b-a55b: cache_read_price_per_mtok $0.10/Mtok to $0.20/Mtok +1 alias | 2026-08-18 |
| truncation | qwen | -50% max out · qwen/qwen3.6-27b: max_output_tokens cut 131,072 to 65,536 | 2026-08-19 |
| cost | qwen | 2.00x price · qwen/qwen3.6-27b: input_price_per_mtok $0.30/Mtok to $0.60/Mtok +1 alias | 2026-08-20 |
| cost | 2.00x price · google/gemma-4-31b-it: cache_read_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-08-21 | |
| truncation | mistral | -50% max out · mistral/codestral-2508: max_output_tokens cut 256,000 to 128,000 | 2026-08-21 |
| truncation | mistral | -50% ctx · mistral/codestral-2508: context_tokens cut 256,000 to 128,000 | 2026-08-21 |
| cost | gemini | 2.00x price · gemini/gemini-3.1-flash-image-preview: input_price_per_mtok $0.25/Mtok to $0.50/Mtok +1 alias | 2026-08-21 |
| cost | gemini | 2.00x price · gemini/gemini-3.1-flash-image-preview: output_price_per_mtok $1.50/Mtok to $3.00/Mtok +1 alias | 2026-08-21 |
| capability | qwen | lost logit_bias, min_p · qwen/qwen3-235b-a22b-thinking-2507: capabilities lost logit_bias, min_p | 2026-08-24 |
| truncation | qwen | -50% ctx · qwen/qwen3-235b-a22b-thinking-2507: context_tokens cut 262,144 to 131,072 | 2026-08-24 |
| cost | deepinfra | 2.00x price · deepinfra/Qwen/Qwen3-14B: input_price_per_mtok $0.06/Mtok to $0.12/Mtok | 2026-08-28 |
| cost | gemini | 2.00x price · gemini/gemini-2.5-pro-preview-tts: output_price_per_mtok $10.00/Mtok to $20.00/Mtok | 2026-08-28 |
| cost | together_ai | 2.00x price · together_ai/Qwen/Qwen3.7-Max: output_price_per_mtok $3.75/Mtok to $7.50/Mtok | 2026-08-28 |
| cost | together_ai | 2.00x price · together_ai/Qwen/Qwen3.7-Max: input_price_per_mtok $1.25/Mtok to $2.50/Mtok | 2026-08-28 |
| cost | vertex_ai-language-models | 2.00x price · gemini-2.5-pro-preview-tts: output_price_per_mtok $10.00/Mtok to $20.00/Mtok | 2026-08-28 |
| truncation | -50% ctx · google/gemma-3-27b-it: context_tokens cut 262,144 to 131,072 | 2026-08-28 | |
| truncation | qwen | -50% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 32,768 to 16,384 +2 aliases | 2026-08-28 |
| cost | 2.00x price · ~google/gemini-flash-latest: cache_read_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-08-28 | |
| cost | 2.00x price · google/gemini-3.7-flash: output_price_per_mtok $1.88/Mtok to $3.75/Mtok | 2026-08-28 | |
| cost | 2.00x price · ~google/gemini-flash-latest: input_price_per_mtok $0.38/Mtok to $0.75/Mtok | 2026-08-28 | |
| cost | 2.00x price · ~google/gemini-flash-latest: output_price_per_mtok $1.88/Mtok to $3.75/Mtok | 2026-08-28 | |
| cost | 2.00x price · google/gemini-3.7-flash: cache_read_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-08-28 | |
| cost | 2.00x price · google/gemini-3.7-flash: input_price_per_mtok $0.38/Mtok to $0.75/Mtok | 2026-08-28 | |
| capability | xai | lost function_calling, tool_choice · xai/grok-4.20-multi-agent-0309: capabilities lost function_calling, tool_choice inferred | 2026-08-30 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.09/Mtok to $0.18/Mtok | 2026-08-30 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +5 aliases | 2026-08-31 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.96/Mtok +5 aliases | 2026-08-31 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.32/Mtok +5 aliases | 2026-08-31 |
| capability | z-ai | lost logprobs, top_logprobs · z-ai/glm-4.7: capabilities lost logprobs, top_logprobs | 2026-08-31 |
| capability | kwaipilot | lost logprobs, top_logprobs · kwaipilot/kat-coder-pro-v2.5: capabilities lost logprobs, top_logprobs +1 alias inferred | 2026-08-31 |
| cost | qwen | 2.00x price · qwen/qwen3.8-2.4t-a95b:batch: cache_read_price_per_mtok $0.25/Mtok to $0.50/Mtok +1 alias | 2026-09-01 |
| capability | qwen | lost logprobs, top_logprobs · qwen/qwen3-32b: capabilities lost logprobs, top_logprobs | 2026-09-01 |
| truncation | gryphe | -50% max out · gryphe/mythomax-l2-13b: max_output_tokens cut 7,372 to 3,686 | 2026-09-01 |
| capability | xiaomi | lost logprobs, top_logprobs · xiaomi/mimo-v2.5: capabilities lost logprobs, top_logprobs | 2026-09-02 |
| truncation | z-ai | -50% max out · z-ai/glm-5.2: max_output_tokens cut 262,144 to 131,072 +7 aliases | 2026-09-02 |
| cost | 2.00x price · google/gemini-3.7-flash:batch: input_price_per_mtok $0.19/Mtok to $0.38/Mtok | 2026-09-02 | |
| cost | 2.00x price · google/gemini-3.7-flash:batch: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok | 2026-09-02 | |
| cost | 2.00x price · google/gemini-3.7-flash:batch: output_price_per_mtok $0.94/Mtok to $1.88/Mtok | 2026-09-02 | |
| cost | qwen | 2.00x price · qwen/qwen3.6-35b-a3b: input_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-09-03 |
| capability | meta-llama | lost logprobs, top_logprobs · meta-llama/llama-3.1-70b-instruct: capabilities lost logprobs, top_logprobs | 2026-09-04 |
| truncation | watsonx | -50% max out · watsonx/bigscience/mt0-xxl-13b: max_output_tokens cut 8,192 to 4,096 inferred | 2026-09-06 |
| truncation | watsonx | -50% ctx · watsonx/bigscience/mt0-xxl-13b: context_tokens cut 8,192 to 4,096 inferred | 2026-09-06 |
| truncation | qwen | -50% max out · qwen/qwen3-14b: max_output_tokens cut 16,384 to 8,192 +3 aliases | 2026-09-08 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-vision-exp: cache_read_price_per_mtok $0.0070/Mtok to $0.01/Mtok +11 aliases | 2026-09-08 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-vision-exp: output_price_per_mtok $0.66/Mtok to $1.32/Mtok +11 aliases | 2026-09-08 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-vision-exp: input_price_per_mtok $0.22/Mtok to $0.44/Mtok +11 aliases | 2026-09-08 |
| truncation | qwen | -50% max out · qwen/qwen3.6-27b: max_output_tokens cut 262,144 to 131,072 +5 aliases | 2026-09-09 |
| capability | z-ai | lost logit_bias, min_p · z-ai/glm-5: capabilities lost logit_bias, min_p | 2026-09-10 |
| capability | thinkingmachines | lost logprobs, top_logprobs · thinkingmachines/inkling-small:free: capabilities lost logprobs, top_logprobs inferred | 2026-09-10 |
| truncation | meta-llama | -50% max out · meta-llama/llama-3.1-70b-instruct: max_output_tokens cut 16,384 to 8,192 +2 aliases | 2026-09-12 |
| cost | openrouter | 2.00x price · openrouter/z-ai/glm-5.3-flash: input_price_per_mtok $0.07/Mtok to $0.15/Mtok | 2026-09-13 |
| cost | openrouter | 2.00x price · openrouter/z-ai/glm-5.3-flash: output_price_per_mtok $0.25/Mtok to $0.50/Mtok | 2026-09-13 |
| cost | openrouter | 2.00x price · openrouter/z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-09-13 |
| cost | z-ai | 2.00x price · z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok +2 aliases | 2026-09-13 |
| cost | z-ai | 2.00x price · z-ai/glm-5.3-flash: output_price_per_mtok $0.25/Mtok to $0.50/Mtok +2 aliases | 2026-09-13 |
| cost | z-ai | 2.00x price · z-ai/glm-5.3-flash: input_price_per_mtok $0.07/Mtok to $0.15/Mtok +2 aliases | 2026-09-13 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.60/Mtok to $1.20/Mtok +3 aliases | 2026-09-16 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0030/Mtok to $0.0060/Mtok +3 aliases | 2026-09-16 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.15/Mtok to $0.30/Mtok +3 aliases | 2026-09-16 |
| truncation | nvidia | -49% ctx · nvidia/nemotron-3-ultra-550b-a55b: context_tokens cut 512,288 to 262,144 | 2026-08-28 |
| truncation | liquid | -49% ctx · liquid/lfm-2.5-2.6b:free: context_tokens cut 128,000 to 65,536 inferred | 2026-08-21 |
| truncation | mistralai | -49% ctx · mistralai/mistral-small-3.2-24b-instruct: context_tokens cut 256,000 to 131,072 | 2026-08-22 |
| truncation | watsonx | -49% max out · watsonx/mistralai/mistral-small-3-1-24b-instruct-2503: max_output_tokens cut 32,000 to 16,384 inferred | 2026-09-06 |
| cost | deepseek | 1.93x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +2 aliases | 2026-09-07 |
| cost | deepseek | 1.93x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.74/Mtok to $3.36/Mtok +2 aliases | 2026-09-07 |
| cost | deepseek | 1.93x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.58/Mtok to $1.12/Mtok +2 aliases | 2026-09-07 |
| cost | nebius | 1.92x price · nebius/Qwen/Qwen2.5-VL-72B-Instruct: input_price_per_mtok $0.13/Mtok to $0.25/Mtok | 2026-09-06 |
| cost | deepseek | 1.90x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +1 alias | 2026-09-11 |
| cost | qwen | 1.90x price · qwen/qwen3-14b: input_price_per_mtok $0.12/Mtok to $0.23/Mtok +1 alias | 2026-09-08 |
| cost | openrouter | 1.90x price · openrouter/qwen/qwen3-14b: input_price_per_mtok $0.12/Mtok to $0.23/Mtok | 2026-09-13 |
| cost | deepseek | 1.89x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.09/Mtok to $0.18/Mtok | 2026-09-14 |
| cost | deepseek | 1.89x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-09-14 |
| cost | deepseek | 1.89x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.0094/Mtok to $0.02/Mtok | 2026-09-14 |
| cost | z-ai | 1.89x price · z-ai/glm-5.2: output_price_per_mtok $1.98/Mtok to $3.74/Mtok | 2026-08-14 |
| cost | z-ai | 1.89x price · z-ai/glm-5.2: input_price_per_mtok $0.63/Mtok to $1.19/Mtok | 2026-08-14 |
| cost | deepseek | 1.89x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.03/Mtok to $0.07/Mtok | 2026-09-11 |
| cost | qwen | 1.88x price · qwen/qwen3.6-27b: input_price_per_mtok $0.32/Mtok to $0.60/Mtok +2 aliases | 2026-08-28 |
| cost | nvidia | 1.88x price · nvidia/nemotron-3-ultra-550b-a55b: cache_read_price_per_mtok $0.10/Mtok to $0.19/Mtok | 2026-09-02 |
| cost | nebius | 1.88x price · nebius/Qwen/Qwen2.5-VL-72B-Instruct: output_price_per_mtok $0.40/Mtok to $0.75/Mtok | 2026-09-06 |
| cost | qwen | 1.87x price · qwen/qwen3-30b-a3b-instruct-2507: input_price_per_mtok $0.05/Mtok to $0.09/Mtok +2 aliases | 2026-09-11 |
| cost | openrouter | 1.87x price · openrouter/qwen/qwen3-30b-a3b-instruct-2507: input_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-09-13 |
| cost | fireworks_ai | 1.86x price · fireworks_ai/accounts/fireworks/models/glm-5p2: cache_read_price_per_mtok $0.14/Mtok to $0.26/Mtok +1 alias | 2026-07-20 |
| cost | z-ai | 1.86x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.14/Mtok to $0.26/Mtok | 2026-09-07 |
| cost | openrouter | 1.85x price · openrouter/mistralai/mistral-small-3.1-24b-instruct: output_price_per_mtok $0.30/Mtok to $0.56/Mtok | 2026-09-06 |
| cost | deepseek | 1.85x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-08-12 |
| cost | deepseek | 1.85x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.63/Mtok to $1.17/Mtok | 2026-08-12 |
| cost | deepseek | 1.85x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.26/Mtok to $2.34/Mtok | 2026-08-12 |
| truncation | ~z-ai | -46% max out · ~z-ai/glm-latest: max_output_tokens cut 235,929 to 128,000 inferred | 2026-09-10 |
| cost | gryphe | 1.83x price · gryphe/mythomax-l2-13b: output_price_per_mtok $0.06/Mtok to $0.11/Mtok | 2026-08-04 |
| cost | qwen | 1.83x price · qwen/qwen3-vl-235b-a22b-instruct: output_price_per_mtok $1.04/Mtok to $1.90/Mtok | 2026-08-17 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.58/Mtok to $1.05/Mtok +3 aliases | 2026-09-14 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.74/Mtok to $3.15/Mtok +3 aliases | 2026-09-14 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-14 |
| cost | deepseek | 1.80x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok | 2026-08-20 |
| cost | deepseek | 1.80x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.56/Mtok | 2026-08-20 |
| cost | qwen | 1.80x price · qwen/qwen3.6-27b: output_price_per_mtok $2.00/Mtok to $3.60/Mtok +1 alias | 2026-08-20 |
| cost | deepseek | 1.80x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.19/Mtok | 2026-08-20 |
| cost | meta-llama | 1.80x price · meta-llama/llama-3.1-70b-instruct: input_price_per_mtok $0.40/Mtok to $0.72/Mtok +1 alias | 2026-09-12 |
| cost | meta-llama | 1.80x price · meta-llama/llama-3.1-70b-instruct: output_price_per_mtok $0.40/Mtok to $0.72/Mtok +1 alias | 2026-09-12 |
| truncation | nvidia | -44% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 235,929 to 131,072 | 2026-08-28 |
| cost | z-ai | 1.79x price · z-ai/glm-5.2: output_price_per_mtok $1.76/Mtok to $3.15/Mtok | 2026-08-12 |
| cost | ~z-ai | 1.79x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.10/Mtok to $0.18/Mtok | 2026-09-03 |
| cost | qwen | 1.79x price · qwen/qwen3.5-35b-a3b: input_price_per_mtok $0.14/Mtok to $0.25/Mtok | 2026-08-13 |
| cost | openrouter | 1.79x price · openrouter/deepseek/deepseek-chat-v3-0324: input_price_per_mtok $0.14/Mtok to $0.25/Mtok | 2026-09-06 |
| cost | deepseek | 1.78x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +1 alias | 2026-09-16 |
| cost | deepseek | 1.78x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0090/Mtok to $0.02/Mtok | 2026-08-30 |
| cost | ~deepseek | 1.78x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.09/Mtok to $0.16/Mtok | 2026-09-07 |
| cost | qwen | 1.76x price · qwen/qwen3.8-27b: cache_read_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-09-12 |
| cost | moonshotai | 1.75x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.28/Mtok to $4.00/Mtok +4 aliases | 2026-08-24 |
| cost | moonshotai | 1.75x price · moonshotai/kimi-k2.6: input_price_per_mtok $0.54/Mtok to $0.95/Mtok +4 aliases | 2026-08-24 |
| cost | moonshotai | 1.75x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.16/Mtok +4 aliases | 2026-08-24 |
| cost | ~deepseek | 1.75x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.14/Mtok to $0.25/Mtok +1 alias | 2026-08-14 |
| cost | deepseek | 1.75x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.08/Mtok to $0.14/Mtok +1 alias | 2026-08-24 |
| cost | ~deepseek | 1.75x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-08-12 |
| cost | deepseek | 1.75x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +1 alias | 2026-08-24 |
| cost | deepseek | 1.74x price · deepseek/deepseek-chat-v3.1: output_price_per_mtok $0.95/Mtok to $1.65/Mtok +1 alias | 2026-09-03 |
| cost | deepinfra | 1.72x price · deepinfra/Qwen/Qwen3-30B-A3B: output_price_per_mtok $0.29/Mtok to $0.50/Mtok | 2026-08-28 |
| cost | 1.71x price · google/gemma-4-26b-a4b-it: input_price_per_mtok $0.07/Mtok to $0.12/Mtok | 2026-08-10 | |
| cost | qwen | 1.70x price · qwen/qwen3.8-27b: cache_read_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-08-25 |
| cost | deepseek | 1.70x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.58/Mtok to $0.98/Mtok +1 alias | 2026-09-16 |
| cost | deepseek | 1.70x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.74/Mtok to $2.95/Mtok +1 alias | 2026-09-16 |
| cost | deepseek | 1.70x price · deepseek/deepseek-v4-pro: output_price_per_mtok $2.34/Mtok to $3.96/Mtok | 2026-08-17 |
| cost | moonshotai | 1.69x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.36/Mtok to $4.00/Mtok +2 aliases | 2026-08-19 |
| cost | moonshotai | 1.69x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.16/Mtok +2 aliases | 2026-08-19 |
| cost | moonshotai | 1.69x price · moonshotai/kimi-k2.6: input_price_per_mtok $0.56/Mtok to $0.95/Mtok +2 aliases | 2026-08-19 |
| cost | deepseek | 1.69x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.12/Mtok +3 aliases | 2026-09-04 |
| cost | deepseek | 1.69x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.35/Mtok +3 aliases | 2026-09-04 |
| cost | deepseek | 1.69x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +3 aliases | 2026-09-04 |
| cost | ~x-ai | 1.67x price · ~x-ai/grok-latest: cache_read_price_per_mtok $0.30/Mtok to $0.50/Mtok | 2026-08-12 |
| cost | gemini | 1.67x price · gemini-2.5-flash-native-audio-latest: input_price_per_mtok $0.30/Mtok to $0.50/Mtok +8 aliases | 2026-08-28 |
| cost | vertex_ai-language-models | 1.67x price · gemini-live-2.5-flash-preview-native-audio-09-2025: input_price_per_mtok $0.30/Mtok to $0.50/Mtok | 2026-08-28 |
| cost | ~deepseek | 1.67x price · ~deepseek/deepseek-v4-flash-latest: input_price_per_mtok $0.03/Mtok to $0.05/Mtok | 2026-08-31 |
| cost | qwen | 1.67x price · qwen/qwen3.5-35b-a3b: output_price_per_mtok $0.75/Mtok to $1.25/Mtok | 2026-09-05 |
| cost | nebius | 1.67x price · nebius/google/gemma-3-27b-it: input_price_per_mtok $0.06/Mtok to $0.10/Mtok | 2026-09-06 |
| cost | ibm-granite | 1.67x price · ibm-granite/granite-4.2-8b: output_price_per_mtok $0.15/Mtok to $0.25/Mtok | 2026-09-09 |
| cost | ~deepseek | 1.67x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0030/Mtok to $0.0050/Mtok | 2026-09-13 |
| cost | qwen | 1.66x price · qwen/qwen3-235b-a22b-2507: input_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-08-03 |
| cost | ~deepseek | 1.66x price · ~deepseek/deepseek-pro-latest: output_price_per_mtok $1.74/Mtok to $2.88/Mtok +1 alias | 2026-09-16 |
| cost | ~deepseek | 1.66x price · ~deepseek/deepseek-pro-latest: input_price_per_mtok $0.58/Mtok to $0.96/Mtok +1 alias | 2026-09-16 |
| cost | deepseek | 1.65x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.63/Mtok to $1.04/Mtok | 2026-09-07 |
| cost | deepseek | 1.65x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.25/Mtok to $2.07/Mtok | 2026-09-07 |
| cost | deepseek | 1.65x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-09-07 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-08-16 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: output_price_per_mtok $2.44/Mtok to $4.00/Mtok +3 aliases | 2026-08-13 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: input_price_per_mtok $0.58/Mtok to $0.95/Mtok +3 aliases | 2026-08-13 |
| cost | moonshotai | 1.64x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.10/Mtok to $0.16/Mtok +3 aliases | 2026-08-13 |
| cost | nvidia | 1.64x price · nvidia/nemotron-3-ultra-550b-a55b: output_price_per_mtok $2.20/Mtok to $3.60/Mtok | 2026-08-01 |
| cost | z-ai | 1.61x price · z-ai/glm-5.2: output_price_per_mtok $1.23/Mtok to $1.98/Mtok | 2026-08-14 |
| cost | z-ai | 1.61x price · z-ai/glm-5.2: input_price_per_mtok $0.39/Mtok to $0.63/Mtok | 2026-08-14 |
| cost | ~deepseek | 1.60x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.10/Mtok to $0.16/Mtok | 2026-09-02 |
| cost | tencent | 1.60x price · tencent/hy3: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +20 aliases | 2026-09-16 |
| cost | tencent | 1.60x price · tencent/hy3: output_price_per_mtok $0.33/Mtok to $0.53/Mtok +20 aliases | 2026-09-16 |
| cost | tencent | 1.60x price · tencent/hy3: input_price_per_mtok $0.08/Mtok to $0.13/Mtok +20 aliases | 2026-09-16 |
| cost | deepinfra | 1.60x price · deepinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning: input_price_per_mtok $0.05/Mtok to $0.08/Mtok | 2026-08-28 |
| cost | openrouter | 1.60x price · openrouter/nvidia/nemotron-3.5-lightning: input_price_per_mtok $0.05/Mtok to $0.08/Mtok | 2026-09-06 |
| cost | qwen | 1.60x price · qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.55/Mtok to $0.88/Mtok | 2026-09-09 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.05/Mtok | 2026-09-08 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.15/Mtok | 2026-09-08 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-08 |
| cost | moonshotai | 1.59x price · moonshotai/kimi-k2.6: cache_read_price_per_mtok $0.09/Mtok to $0.15/Mtok | 2026-08-15 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.18/Mtok to $0.28/Mtok | 2026-08-07 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-07 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.09/Mtok to $0.14/Mtok | 2026-08-07 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.06/Mtok to $0.09/Mtok | 2026-08-25 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.11/Mtok to $0.18/Mtok | 2026-08-25 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-08-25 |
| cost | ~deepseek | 1.58x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-11 |
| cost | ~deepseek | 1.58x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.16/Mtok to $0.25/Mtok | 2026-08-11 |
| cost | z-ai | 1.58x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.14/Mtok to $0.22/Mtok | 2026-08-17 |
| cost | ~deepseek | 1.57x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-08-14 |
| cost | deepseek | 1.57x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.09/Mtok to $0.14/Mtok | 2026-09-01 |
| cost | qwen | 1.57x price · qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.35/Mtok to $0.55/Mtok +1 alias | 2026-09-05 |
| cost | fireworks_ai | 1.57x price · fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_price_per_mtok $0.14/Mtok to $0.22/Mtok +1 alias | 2026-09-01 |
| cost | z-ai | 1.57x price · z-ai/glm-5.2: input_price_per_mtok $0.76/Mtok to $1.19/Mtok | 2026-08-17 |
| cost | nvidia | 1.56x price · nvidia/nemotron-3-ultra-550b-a55b: cache_read_price_per_mtok $0.12/Mtok to $0.19/Mtok +1 alias | 2026-09-16 |
| cost | deepseek | 1.56x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.18/Mtok to $0.28/Mtok +1 alias | 2026-08-24 |
Announced deprecations
Changes that came with a published date. Separated out because they are a different problem: you were told.
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| availability | friendliai | friendliai/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B: retires 2026-09-06 (7 days ago) inferred | 2026-09-13 |
| availability | together_ai | together_ai/moonshotai/Kimi-K2.6: retires 2026-08-19 (25 days ago) | 2026-09-13 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: retires 2026-09-14 (in 1 days) +3 aliases | 2026-09-13 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: deprecation announced for 2026-09-14 — in 1 days +3 aliases | 2026-09-13 |
| availability | openai | gpt-5.4-cyber: retires 2026-10-01 (in 19 days) inferred | 2026-09-12 |
| availability | dots-studio | dots-studio/dots-3-note-preview:free: retirement announced for 2026-09-30 — in 19 days +1 alias inferred | 2026-09-11 |
| availability | scaleway | scaleway/mistralai/pixtral-12b-2409: retires 2026-10-01 (in 21 days) +1 alias | 2026-09-10 |
| availability | scaleway | scaleway/mistralai/pixtral-12b-2409: deprecation announced for 2026-10-01 — in 21 days +1 alias | 2026-09-10 |
| availability | bedrock_converse | us-gov.anthropic.claude-3-haiku-20240307-v1:0: retires 2026-09-10 (in 1 days) inferred | 2026-09-09 |
| availability | azure | azure/gpt-realtime-2: retires 2026-08-31 (6 days ago) inferred | 2026-09-06 |
| availability | databricks | databricks/databricks-gemini-2-5-flash: deprecation announced for 2026-10-02 — in 26 days inferred | 2026-09-06 |
| availability | databricks | databricks/databricks-gemini-2-5-flash: retires 2026-10-02 (in 26 days) inferred | 2026-09-06 |
| availability | z-ai | z-ai/glm-4.7-flash: retirement announced for 2026-09-10 — in 6 days | 2026-09-04 |
| availability | z-ai | z-ai/glm-4.7-flash: retires 2026-09-10 (in 6 days) | 2026-09-04 |
| availability | azure_ai | azure_ai/MAI-Image-2.5-Flash: retires 2026-10-01 (in 28 days) +1 alias inferred | 2026-09-03 |
| availability | azure_ai | azure_ai/MAI-Image-2.5-Flash: deprecation announced for 2026-10-01 — in 28 days +1 alias inferred | 2026-09-03 |
| availability | cerebras | cerebras/zai-glm-4.7: retires 2026-08-17 (15 days ago) inferred | 2026-09-01 |
| availability | bedrock | amazon.nova-sonic-v1:0: retires 2026-09-14 (in 13 days) inferred | 2026-09-01 |
| availability | cerebras | cerebras/zai-glm-4.7: deprecation announced for 2026-08-17 (date already passed) inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-mini: retires 2026-09-08 (in 7 days) +1 alias inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-mini: retirement announced for 2026-09-08 — in 7 days +1 alias inferred | 2026-09-01 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Pro: deprecation announced for 2026-08-27 (date already passed) +3 aliases | 2026-08-28 |
| availability | together_ai | together_ai/google/gemma-3n-E4B-it: deprecation announced for 2026-08-25 (date already passed) +1 alias | 2026-08-28 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Pro: retires 2026-08-27 (1 days ago) +3 aliases | 2026-08-28 |
| availability | together_ai | together_ai/google/gemma-3n-E4B-it: retires 2026-08-25 (3 days ago) +1 alias | 2026-08-28 |
| availability | moonshotai | moonshotai/kimi-k2.5: retirement announced for 2026-08-31 — in 4 days +1 alias | 2026-08-27 |
| availability | moonshotai | moonshotai/kimi-k2.5: retires 2026-08-31 (in 7 days) | 2026-08-24 |
| availability | azure_ai | azure_ai/deepseek-r1: deprecation announced for 2026-08-13 (date already passed) | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) inferred | 2026-08-21 |
| availability | gemini | gemini/gemini-robotics-er-1.6-preview: retires 2026-08-31 (in 10 days) | 2026-08-21 |
| availability | azure_ai | azure_ai/deepseek-r1: retires 2026-08-13 (8 days ago) | 2026-08-21 |
| availability | azure_ai | azure_ai/MAI-Image-2e: deprecation announced for 2026-08-15 (date already passed) inferred | 2026-08-21 |
| availability | gemini | gemini/gemini-robotics-er-1.6-preview: deprecation announced for 2026-08-31 — in 10 days | 2026-08-21 |
| availability | azure_ai | azure_ai/MAI-Image-2e: retires 2026-08-15 (6 days ago) inferred | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-opus-4-1: retires 2026-08-05 (16 days ago) inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) +1 alias inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-opus-4-1: retires 2026-08-05 (16 days ago) +1 alias inferred | 2026-08-21 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free: retires 2026-08-24 (in 4 days) +2 aliases inferred | 2026-08-20 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free: retirement announced for 2026-08-24 — in 4 days +2 aliases inferred | 2026-08-20 |
| availability | inclusionai | inclusionai/ling-2.6-1t: retires 2026-08-24 (in 5 days) +2 aliases inferred | 2026-08-19 |
| availability | inclusionai | inclusionai/ling-2.6-1t: retirement announced for 2026-08-24 — in 5 days +2 aliases inferred | 2026-08-19 |
| availability | deepseek | deepseek/deepseek-v3.1-terminus: retires 2026-08-17 (in 2 days) | 2026-08-15 |
| availability | deepseek | deepseek/deepseek-v3.1-terminus: retirement announced for 2026-08-17 — in 2 days | 2026-08-15 |
| availability | groq | groq/llama-3.1-8b-instant: deprecation announced for 2026-08-16 — in 3 days +1 alias inferred | 2026-08-13 |
| availability | groq | groq/meta-llama/llama-4-scout-17b-16e-instruct: deprecation announced for 2026-07-17 (date already passed) +1 alias | 2026-08-13 |
| availability | groq | groq/meta-llama/llama-4-scout-17b-16e-instruct: retires 2026-07-17 (27 days ago) +1 alias | 2026-08-13 |
| availability | groq | groq/llama-3.1-8b-instant: retires 2026-08-16 (in 3 days) +1 alias inferred | 2026-08-13 |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free: retirement announced for 2026-08-13 — in 1 days inferred | 2026-08-12 |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free: retires 2026-08-13 (in 1 days) inferred | 2026-08-12 |
| availability | mistral | mistral/devstral-2512: retires 2026-07-31 (12 days ago) +5 aliases inferred | 2026-08-12 |
| availability | anthropic | claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-sonnet-20240229-v1:0: retires 2026-07-30 (13 days ago) +8 aliases | 2026-08-12 |
| availability | gemini | gemini/imagen-4.0-fast-generate-001: retires 2026-08-17 (in 5 days) +2 aliases | 2026-08-12 |
| availability | bedrock | cohere.command-r-plus-v1:0: retires 2026-08-19 (in 7 days) +1 alias inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-haiku-20240307-v1:0: deprecation announced for 2026-09-10 — in 29 days +5 aliases | 2026-08-12 |
| availability | mistral | mistral/mistral-medium-2505: deprecation announced for 2026-08-31 — in 19 days +2 aliases inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-haiku-20240307-v1:0: retires 2026-09-10 (in 29 days) +5 aliases | 2026-08-12 |
| availability | gemini | gemini/imagen-4.0-fast-generate-001: deprecation announced for 2026-08-17 — in 5 days +2 aliases | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-sonnet-20240229-v1:0: deprecation announced for 2026-07-30 (date already passed) +8 aliases | 2026-08-12 |
| availability | mistral | mistral/mistral-medium-2505: retires 2026-08-31 (in 19 days) +2 aliases inferred | 2026-08-12 |
| availability | bedrock | cohere.command-r-plus-v1:0: deprecation announced for 2026-08-19 — in 7 days +1 alias inferred | 2026-08-12 |
| availability | anthropic | claude-opus-4-1-20250805: retires 2026-08-05 (in 4 days) +1 alias inferred | 2026-08-12 |
| availability | mistral | mistral/devstral-2512: deprecation announced for 2026-07-31 (date already passed) +5 aliases inferred | 2026-08-12 |
| availability | gemini | gemini/gemini-embedding-2-preview: retires 2026-08-10 (2 days ago) | 2026-08-12 |
| availability | gemini | gemini/gemini-embedding-2-preview: deprecation announced for 2026-08-10 (date already passed) | 2026-08-12 |
| availability | openai | gpt-5.2-chat-latest: deprecation announced for 2026-08-10 (date already passed) +1 alias | 2026-08-12 |
| availability | openai | gpt-4o-mini-search-preview-2025-03-11: deprecation announced for 2026-07-23 (date already passed) +15 aliases | 2026-08-12 |
| availability | openai | gpt-4o-audio: deprecation announced for 2026-07-20 (date already passed) +8 aliases inferred | 2026-08-05 |
| availability | openai | computer-use-preview-2025-03-11: retires 2026-07-23 (9 days ago) +17 aliases inferred | 2026-08-01 |
| availability | openai | gpt-5.2-chat-latest: retires 2026-08-10 (in 9 days) +3 aliases | 2026-08-01 |
| availability | azure | azure/chat-latest: retires 2026-12-02 (in 79 days) +2 aliases | 2026-09-14 |
| availability | xai | xai/grok-imagine-image-quality-20260403: retires 2026-11-02 (in 50 days) +2 aliases inferred | 2026-09-13 |
| availability | azure | azure/eu/o1-2024-12-17: retires 2026-11-19 (in 67 days) +16 aliases | 2026-09-13 |
| availability | azure | azure/eu/o1-2024-12-17: deprecation announced for 2026-11-19 — in 67 days +6 aliases | 2026-09-13 |
| availability | azure | azure/o4-mini-2025-04-16: deprecation announced for 2026-11-19 — in 67 days +2 aliases | 2026-09-13 |
| availability | azure | azure/eu/o3-mini-2025-01-31: deprecation announced for 2026-11-19 — in 67 days +3 aliases | 2026-09-13 |
| availability | azure | azure/o3-deep-research: deprecation announced for 2026-11-19 — in 67 days | 2026-09-13 |
| availability | azure | azure/o3-pro-2025-06-10: deprecation announced for 2026-11-19 — in 67 days +1 alias | 2026-09-13 |
| availability | xai | xai/grok-imagine-image-quality-20260403: deprecation announced for 2026-11-02 — in 50 days +2 aliases inferred | 2026-09-13 |
| availability | azure_ai | azure_ai/gpt-chat-latest: retires 2026-12-02 (in 83 days) | 2026-09-10 |
| availability | azure_ai | azure_ai/codex-mini: retires 2026-11-15 (in 66 days) inferred | 2026-09-10 |
| availability | databricks | databricks/databricks-claude-sonnet-4: retires 2026-10-09 (in 33 days) inferred | 2026-09-06 |
| availability | databricks | databricks/databricks-claude-sonnet-4: deprecation announced for 2026-10-09 — in 33 days inferred | 2026-09-06 |
| availability | azure_ai | azure_ai/kimi-k2.7-code: retires 2026-10-03 (in 30 days) | 2026-09-03 |
| availability | gemini | gemini/gemini-omni-flash-preview: deprecation announced for 2026-09-30 — in 31 days | 2026-08-30 |
| availability | gemini | gemini/gemini-omni-flash-preview: retires 2026-09-30 (in 31 days) | 2026-08-30 |
| availability | azure | azure/gpt-4o-transcribe: retires 2026-10-15 (in 64 days) +1 alias inferred | 2026-08-28 |
| availability | azure | azure/gpt-4o-transcribe: deprecation announced for 2026-10-15 — in 64 days +1 alias inferred | 2026-08-28 |
| availability | anthropic | claude-sonnet-4-5-20250929: retires 2026-09-29 (in 39 days) +1 alias | 2026-08-21 |
| availability | azure | azure/eu/o3-mini-2025-01-31: retires 2026-10-01 (in 50 days) +5 aliases | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-lite: retires 2026-10-20 (in 60 days) +2 aliases | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-haiku-4-5: retires 2026-10-19 (in 59 days) +2 aliases inferred | 2026-08-21 |
| availability | azure | azure/gpt-4.1-nano-2025-04-14: retires 2026-10-14 (in 63 days) +2 aliases | 2026-08-21 |
| availability | text-completion-openai | babbage-002: retires 2026-09-28 (in 38 days) +2 aliases | 2026-08-21 |
| availability | azure | azure/o4-mini-2025-04-16: deprecation announced for 2026-10-16 — in 65 days +2 aliases | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-lite: deprecation announced for 2026-10-20 — in 60 days +2 aliases | 2026-08-21 |
| availability | azure | azure/gpt-4.1-nano: deprecation announced for 2026-10-14 — in 54 days | 2026-08-21 |
| availability | vertex_ai-video-models | vertex_ai/veo-3.1-fast-generate-001: deprecation announced for 2026-11-17 — in 88 days +1 alias inferred | 2026-08-21 |
| availability | anthropic | claude-haiku-4-5-20251001: deprecation announced for 2026-10-15 — in 55 days +1 alias | 2026-08-21 |
| availability | openai | ft-babbage-002: retires 2026-10-23 (in 83 days) +45 aliases | 2026-08-21 |
| availability | azure | azure/o4-mini-2025-04-16: retires 2026-10-16 (in 65 days) +2 aliases | 2026-08-21 |
| availability | azure | azure/gpt-image-1: retires 2026-10-23 (in 72 days) +9 aliases | 2026-08-21 |
| availability | azure | azure/eu/o1-2024-12-17: retires 2026-10-21 (in 70 days) +6 aliases | 2026-08-21 |
| availability | anthropic | claude-haiku-4-5-20251001: retires 2026-10-15 (in 55 days) +1 alias | 2026-08-21 |
| availability | openai | ft:gpt-3.5-turbo-0125: deprecation announced for 2026-10-23 — in 72 days +33 aliases | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-sonnet-4-5: deprecation announced for 2026-09-29 — in 39 days +1 alias inferred | 2026-08-21 |
| availability | azure | azure/eu/o3-mini-2025-01-31: deprecation announced for 2026-10-01 — in 50 days +4 aliases | 2026-08-21 |
| availability | anthropic | claude-sonnet-4-5-20250929: deprecation announced for 2026-09-29 — in 39 days +1 alias | 2026-08-21 |
| availability | vertex_ai-video-models | vertex_ai/veo-3.1-fast-generate-001: retires 2026-11-17 (in 88 days) +1 alias inferred | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-image: deprecation announced for 2026-10-02 — in 42 days +1 alias | 2026-08-21 |
| availability | azure | azure/eu/o1-2024-12-17: deprecation announced for 2026-10-21 — in 70 days +4 aliases | 2026-08-21 |
| availability | vertex_ai-language-models | gemini-2.5-flash-image: retires 2026-10-02 (in 42 days) +1 alias | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-haiku-4-5: deprecation announced for 2026-10-19 — in 59 days +2 aliases inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-haiku-4-5: deprecation announced for 2026-10-15 — in 55 days +1 alias inferred | 2026-08-21 |
| availability | text-completion-openai | babbage-002: deprecation announced for 2026-09-28 — in 38 days +2 aliases | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-haiku-4-5: retires 2026-10-15 (in 55 days) +1 alias inferred | 2026-08-21 |
| availability | azure | azure/gpt-image-1: deprecation announced for 2026-10-23 — in 72 days +9 aliases | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-sonnet-4-5: retires 2026-09-29 (in 39 days) +1 alias inferred | 2026-08-21 |
| availability | dots-studio | dots-studio/dots-3-note-preview:free: retires 2026-09-30 (in 41 days) inferred | 2026-08-20 |
| availability | azure | azure/gpt-4.1-nano-2025-04-14: deprecation announced for 2026-10-14 — in 63 days +1 alias | 2026-08-12 |
| availability | text-completion-openai | ft:babbage-002: retires 2026-10-23 (in 72 days) +1 alias inferred | 2026-08-12 |
| availability | openai | gpt-4-1106-preview: deprecation announced for 2026-10-23 — in 72 days | 2026-08-12 |
| availability | bedrock_converse | anthropic.claude-sonnet-4-20250514-v1:0: retires 2026-10-14 (in 63 days) +4 aliases | 2026-08-12 |
| availability | bedrock_converse | anthropic.claude-sonnet-4-20250514-v1:0: deprecation announced for 2026-10-14 — in 63 days +4 aliases | 2026-08-12 |
| availability | bedrock_converse | us.amazon.nova-premier-v1:0: deprecation announced for 2026-09-14 — in 33 days inferred | 2026-08-12 |
| availability | azure | azure/o3-2025-04-16: deprecation announced for 2026-10-21 — in 70 days +1 alias | 2026-08-12 |
| availability | bedrock_converse | us.amazon.nova-premier-v1:0: retires 2026-09-14 (in 33 days) inferred | 2026-08-12 |
| availability | gemini | gemini/gemini-2.5-flash-image: deprecation announced for 2026-10-02 — in 51 days | 2026-08-12 |
| availability | bedrock | 1024-x-1024/50-steps/bedrock/amazon.nova-canvas-v1:0: deprecation announced for 2026-09-30 — in 49 days +2 aliases inferred | 2026-08-12 |
| availability | openai | gpt-4-0613: deprecation announced for 2026-10-23 — in 72 days | 2026-08-12 |
| availability | gemini | gemini/gemini-2.5-flash-image: retires 2026-10-02 (in 51 days) | 2026-08-12 |
| availability | text-completion-openai | ft:babbage-002: deprecation announced for 2026-10-23 — in 72 days +1 alias inferred | 2026-08-12 |
| availability | bedrock | 1024-x-1024/50-steps/bedrock/amazon.nova-canvas-v1:0: retires 2026-09-30 (in 49 days) +2 aliases inferred | 2026-08-12 |
| availability | openai | openai/sora-2-pro: deprecation announced for 2026-09-24 — in 43 days +3 aliases | 2026-08-12 |
| availability | openai | openai/sora-2-pro: retires 2026-09-24 (in 43 days) +6 aliases | 2026-08-12 |