AI API changes: September 2026
195 changes shipped without an announcement, and 51 that came with a published date. Compiled automatically from provider catalogs and endpoints; every figure below links to the model it came from.
44
capacity cuts
context or output shrank
115
price changes
same call, different bill
11
capabilities lost
51
announced
deprecations with a date
Unannounced, biggest first
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| truncation | vertex_ai | -94% ctx · vertex_ai/xai/grok-4.1-fast-non-reasoning: context_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| truncation | vertex_ai | -94% max out · vertex_ai/xai/grok-4.1-fast-non-reasoning: max_output_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| truncation | watsonx | -94% max out · watsonx/meta-llama/llama-4-maverick-17b: max_output_tokens cut 128,000 to 8,192 inferred | 2026-09-06 |
| truncation | openrouter | -94% max out · openrouter/anthropic/claude-sonnet-4.5: max_output_tokens cut 1,000,000 to 64,000 | 2026-09-06 |
| truncation | thinkingmachines | -93% max out · thinkingmachines/inkling: max_output_tokens cut 471,859 to 32,768 | 2026-09-10 |
| truncation | qwen | -93% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 235,929 to 16,384 +5 aliases | 2026-09-09 |
| cost | gryphe | 10.00x price · gryphe/mythomax-l2-13b: output_price_per_mtok $0.06/Mtok to $0.60/Mtok | 2026-09-01 |
| cost | openrouter | 9.23x price · openrouter/mistralai/mixtral-8x22b-instruct: output_price_per_mtok $0.65/Mtok to $6.00/Mtok | 2026-09-06 |
| truncation | mistralai | -88% max out · mistralai/mistral-medium-3-5:batch: max_output_tokens cut 209,715 to 26,214 inferred | 2026-09-04 |
| truncation | vertex_ai-language-models | -88% ctx · gemini-live-2.5-flash-native-audio: context_tokens cut 1,048,576 to 131,072 inferred | 2026-09-10 |
| truncation | z-ai | -88% max out · z-ai/glm-4.6: max_output_tokens cut 131,072 to 16,384 +2 aliases | 2026-09-09 |
| truncation | mistralai | -88% ctx · mistralai/mistral-medium-3-5:batch: context_tokens cut 262,144 to 32,768 inferred | 2026-09-04 |
| truncation | ~z-ai | -86% max out · ~z-ai/glm-latest: max_output_tokens cut 943,718 to 128,000 inferred | 2026-09-08 |
| truncation | qwen | -86% max out · qwen/qwen3-30b-a3b-instruct-2507: max_output_tokens cut 235,929 to 32,000 +1 alias | 2026-09-10 |
| truncation | ~z-ai | -86% max out · ~z-ai/glm-flash-latest: max_output_tokens cut 943,718 to 131,072 +3 aliases inferred | 2026-09-09 |
| truncation | deepseek | -86% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 943,718 to 131,072 | 2026-09-06 |
| truncation | meta-llama | -86% max out · meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 115,200 to 16,384 | 2026-09-03 |
| cost | gryphe | 6.67x price · gryphe/mythomax-l2-13b: input_price_per_mtok $0.06/Mtok to $0.40/Mtok | 2026-09-01 |
| truncation | nvidia | -82% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 182,520 to 32,768 | 2026-09-04 |
| cost | openrouter | 5.56x price · openrouter/qwen/qwen-2.5-coder-32b-instruct: output_price_per_mtok $0.18/Mtok to $1.00/Mtok | 2026-09-06 |
| truncation | openrouter | -80% ctx · openrouter/anthropic/claude-sonnet-4.5: context_tokens cut 1,000,000 to 200,000 | 2026-09-06 |
| truncation | deepseek | -77% max out · deepseek/deepseek-chat-v3.1: max_output_tokens cut 144,900 to 32,768 +1 alias | 2026-09-08 |
| cost | deepseek | 4.23x price · deepseek/deepseek-chat-v3.1: cache_read_price_per_mtok $0.13/Mtok to $0.55/Mtok +1 alias | 2026-09-03 |
| truncation | nebius | -76% ctx · nebius/Qwen/Qwen2.5-VL-72B-Instruct: context_tokens cut 131,072 to 32,000 | 2026-09-06 |
| truncation | nebius | -76% max out · nebius/Qwen/Qwen2.5-VL-72B-Instruct: max_output_tokens cut 131,072 to 32,000 | 2026-09-06 |
| truncation | ~z-ai | -75% max out · ~z-ai/glm-latest: max_output_tokens cut 943,718 to 235,929 +1 alias inferred | 2026-09-04 |
| capability | vertex_ai-language-models | lost pdf_input, prompt_caching, response_schema, url_context · gemini-live-2.5-flash-native-audio: capabilities lost pdf_input, prompt_caching, response_schema, url_context inferred | 2026-09-10 |
| truncation | openai | -75% ctx · gpt-realtime-mini: context_tokens cut 128,000 to 32,000 | 2026-09-03 |
| cost | qwen | 3.91x price · qwen/qwen3.5-35b-a3b: input_price_per_mtok $0.08/Mtok to $0.31/Mtok | 2026-09-05 |
| cost | openrouter | 3.83x price · openrouter/qwen/qwen3-235b-a22b-thinking-2507: output_price_per_mtok $0.60/Mtok to $2.30/Mtok | 2026-09-06 |
| truncation | nvidia | -74% ctx · nvidia/nemotron-3-super-120b-a12b: context_tokens cut 1,000,000 to 262,144 +1 alias | 2026-09-09 |
| cost | qwen | 3.79x price · qwen/qwen3-14b: output_price_per_mtok $0.24/Mtok to $0.91/Mtok +1 alias | 2026-09-08 |
| cost | z-ai | 3.72x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.05/Mtok to $0.19/Mtok | 2026-09-10 |
| cost | openrouter | 3.67x price · openrouter/qwen/qwen-2.5-coder-32b-instruct: input_price_per_mtok $0.18/Mtok to $0.66/Mtok | 2026-09-06 |
| truncation | qwen | -72% max out · qwen/qwen3.5-35b-a3b: max_output_tokens cut 235,929 to 65,536 +3 aliases | 2026-09-06 |
| cost | openrouter | 3.57x price · openrouter/deepseek/deepseek-chat-v3-0324: output_price_per_mtok $0.28/Mtok to $1.00/Mtok | 2026-09-06 |
| cost | openrouter | 3.51x price · openrouter/mistralai/mistral-small-3.1-24b-instruct: input_price_per_mtok $0.10/Mtok to $0.35/Mtok | 2026-09-06 |
| cost | openrouter | 3.50x price · openrouter/qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.10/Mtok to $0.35/Mtok | 2026-09-06 |
| cost | z-ai | 3.45x price · z-ai/glm-5.2: output_price_per_mtok $0.88/Mtok to $3.04/Mtok | 2026-09-10 |
| cost | z-ai | 3.45x price · z-ai/glm-5.2: input_price_per_mtok $0.28/Mtok to $0.97/Mtok | 2026-09-10 |
| cost | openrouter | 3.33x price · openrouter/mistralai/devstral-2512: output_price_per_mtok $0.60/Mtok to $2.00/Mtok | 2026-09-06 |
| cost | bedrock | 3.20x price · eu.anthropic.claude-3-5-haiku-20241022-v1:0: input_price_per_mtok $0.25/Mtok to $0.80/Mtok | 2026-09-10 |
| cost | bedrock | 3.20x price · eu.anthropic.claude-3-5-haiku-20241022-v1:0: output_price_per_mtok $1.25/Mtok to $4.00/Mtok | 2026-09-10 |
| cost | qwen | 3.20x price · qwen/qwen2.5-vl-72b-instruct: input_price_per_mtok $0.25/Mtok to $0.80/Mtok +1 alias | 2026-09-03 |
| cost | bedrock | 3.20x price · eu.anthropic.claude-3-5-haiku-20241022-v1:0: cache_read_price_per_mtok $0.03/Mtok to $0.08/Mtok | 2026-09-10 |
| cost | openrouter | 3.18x price · openrouter/deepseek/deepseek-chat: output_price_per_mtok $0.28/Mtok to $0.89/Mtok | 2026-09-06 |
| truncation | openrouter | -68% max out · openrouter/anthropic/claude-haiku-4.5: max_output_tokens cut 200,000 to 64,000 | 2026-09-06 |
| cost | openrouter | 3.08x price · openrouter/mistralai/mixtral-8x22b-instruct: input_price_per_mtok $0.65/Mtok to $2.00/Mtok | 2026-09-06 |
| capability | nvidia | lost logprobs, structured_outputs, top_logprobs · nvidia/nemotron-3-super-120b-a12b: capabilities lost logprobs, structured_outputs, top_logprobs | 2026-09-09 |
| truncation | ~deepseek | -67% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 393,216 to 131,072 +1 alias inferred | 2026-09-04 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.10/Mtok to $0.28/Mtok | 2026-09-07 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0100/Mtok to $0.03/Mtok | 2026-09-07 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.05/Mtok to $0.14/Mtok | 2026-09-07 |
| cost | openrouter | 2.67x price · openrouter/mistralai/devstral-2512: input_price_per_mtok $0.15/Mtok to $0.40/Mtok | 2026-09-06 |
| cost | deepseek | 2.63x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.06/Mtok | 2026-09-08 |
| cost | qwen | 2.44x price · qwen/qwen3-235b-a22b-2507: input_price_per_mtok $0.09/Mtok to $0.22/Mtok | 2026-09-09 |
| truncation | ~deepseek | -58% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 943,718 to 393,216 inferred | 2026-09-07 |
| cost | fireworks_ai | 2.36x price · fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: output_price_per_mtok $0.28/Mtok to $0.66/Mtok +1 alias | 2026-09-01 |
| truncation | moonshotai | -57% max out · moonshotai/kimi-k2-thinking: max_output_tokens cut 235,929 to 100,352 +1 alias | 2026-09-10 |
| cost | ~z-ai | 2.34x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.10/Mtok to $0.23/Mtok | 2026-09-06 |
| cost | openrouter | 2.29x price · openrouter/deepseek/deepseek-chat: input_price_per_mtok $0.14/Mtok to $0.32/Mtok | 2026-09-06 |
| cost | deepseek | 2.20x price · deepseek/deepseek-chat-v3.1: input_price_per_mtok $0.25/Mtok to $0.55/Mtok +1 alias | 2026-09-03 |
| cost | openrouter | 2.09x price · openrouter/qwen/qwen3-235b-a22b-thinking-2507: input_price_per_mtok $0.11/Mtok to $0.23/Mtok | 2026-09-06 |
| cost | bedrock_converse | 2.05x price · qwen.qwen3-coder-480b-a35b-v1:0: input_price_per_mtok $0.22/Mtok to $0.45/Mtok | 2026-09-06 |
| capability | z-ai | lost logit_bias, min_p · z-ai/glm-5: capabilities lost logit_bias, min_p | 2026-09-10 |
| truncation | qwen | -50% max out · qwen/qwen3.6-27b: max_output_tokens cut 262,144 to 131,072 +5 aliases | 2026-09-09 |
| truncation | qwen | -50% max out · qwen/qwen3-14b: max_output_tokens cut 16,384 to 8,192 +3 aliases | 2026-09-08 |
| truncation | watsonx | -50% max out · watsonx/bigscience/mt0-xxl-13b: max_output_tokens cut 8,192 to 4,096 inferred | 2026-09-06 |
| truncation | watsonx | -50% ctx · watsonx/bigscience/mt0-xxl-13b: context_tokens cut 8,192 to 4,096 inferred | 2026-09-06 |
| capability | meta-llama | lost logprobs, top_logprobs · meta-llama/llama-3.1-70b-instruct: capabilities lost logprobs, top_logprobs | 2026-09-04 |
| capability | xiaomi | lost logprobs, top_logprobs · xiaomi/mimo-v2.5: capabilities lost logprobs, top_logprobs | 2026-09-02 |
| truncation | z-ai | -50% max out · z-ai/glm-5.2: max_output_tokens cut 262,144 to 131,072 +7 aliases | 2026-09-02 |
| capability | qwen | lost logprobs, top_logprobs · qwen/qwen3-32b: capabilities lost logprobs, top_logprobs | 2026-09-01 |
| truncation | gryphe | -50% max out · gryphe/mythomax-l2-13b: max_output_tokens cut 7,372 to 3,686 | 2026-09-01 |
| cost | z-ai | 2.00x price · z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-09-10 |
| cost | z-ai | 2.00x price · z-ai/glm-5.3-flash: output_price_per_mtok $0.25/Mtok to $0.50/Mtok | 2026-09-10 |
| cost | z-ai | 2.00x price · z-ai/glm-5.3-flash: input_price_per_mtok $0.07/Mtok to $0.15/Mtok | 2026-09-10 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-vision-exp: cache_read_price_per_mtok $0.0070/Mtok to $0.01/Mtok +11 aliases | 2026-09-08 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-vision-exp: output_price_per_mtok $0.66/Mtok to $1.32/Mtok +11 aliases | 2026-09-08 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-flash-vision-exp: input_price_per_mtok $0.22/Mtok to $0.44/Mtok +11 aliases | 2026-09-08 |
| cost | qwen | 2.00x price · qwen/qwen3.6-35b-a3b: input_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-09-03 |
| cost | 2.00x price · google/gemini-3.7-flash:batch: input_price_per_mtok $0.19/Mtok to $0.38/Mtok | 2026-09-02 | |
| cost | 2.00x price · google/gemini-3.7-flash:batch: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok | 2026-09-02 | |
| cost | 2.00x price · google/gemini-3.7-flash:batch: output_price_per_mtok $0.94/Mtok to $1.88/Mtok | 2026-09-02 | |
| cost | qwen | 2.00x price · qwen/qwen3.8-2.4t-a95b:batch: cache_read_price_per_mtok $0.25/Mtok to $0.50/Mtok +1 alias | 2026-09-01 |
| truncation | watsonx | -49% max out · watsonx/mistralai/mistral-small-3-1-24b-instruct-2503: max_output_tokens cut 32,000 to 16,384 inferred | 2026-09-06 |
| cost | deepseek | 1.93x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +2 aliases | 2026-09-07 |
| cost | deepseek | 1.93x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.74/Mtok to $3.36/Mtok +2 aliases | 2026-09-07 |
| cost | deepseek | 1.93x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.58/Mtok to $1.12/Mtok +2 aliases | 2026-09-07 |
| cost | nebius | 1.92x price · nebius/Qwen/Qwen2.5-VL-72B-Instruct: input_price_per_mtok $0.13/Mtok to $0.25/Mtok | 2026-09-06 |
| cost | deepseek | 1.90x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-10 |
| cost | qwen | 1.90x price · qwen/qwen3-14b: input_price_per_mtok $0.12/Mtok to $0.23/Mtok +1 alias | 2026-09-08 |
| cost | nebius | 1.88x price · nebius/Qwen/Qwen2.5-VL-72B-Instruct: output_price_per_mtok $0.40/Mtok to $0.75/Mtok | 2026-09-06 |
| cost | nvidia | 1.88x price · nvidia/nemotron-3-ultra-550b-a55b: cache_read_price_per_mtok $0.10/Mtok to $0.19/Mtok | 2026-09-02 |
| cost | qwen | 1.87x price · qwen/qwen3-30b-a3b-instruct-2507: input_price_per_mtok $0.05/Mtok to $0.09/Mtok +1 alias | 2026-09-10 |
| cost | z-ai | 1.86x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.14/Mtok to $0.26/Mtok | 2026-09-07 |
| cost | openrouter | 1.85x price · openrouter/mistralai/mistral-small-3.1-24b-instruct: output_price_per_mtok $0.30/Mtok to $0.56/Mtok | 2026-09-06 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.58/Mtok to $1.05/Mtok +1 alias | 2026-09-10 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.74/Mtok to $3.15/Mtok +1 alias | 2026-09-10 |
| cost | ~z-ai | 1.79x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.10/Mtok to $0.18/Mtok | 2026-09-03 |
| cost | openrouter | 1.79x price · openrouter/deepseek/deepseek-chat-v3-0324: input_price_per_mtok $0.14/Mtok to $0.25/Mtok | 2026-09-06 |
| cost | ~deepseek | 1.78x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.09/Mtok to $0.16/Mtok | 2026-09-07 |
| cost | deepseek | 1.74x price · deepseek/deepseek-chat-v3.1: output_price_per_mtok $0.95/Mtok to $1.65/Mtok +1 alias | 2026-09-03 |
| cost | deepseek | 1.69x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.12/Mtok +3 aliases | 2026-09-04 |
| cost | deepseek | 1.69x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.35/Mtok +3 aliases | 2026-09-04 |
| cost | deepseek | 1.69x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +3 aliases | 2026-09-04 |
| cost | ibm-granite | 1.67x price · ibm-granite/granite-4.2-8b: output_price_per_mtok $0.15/Mtok to $0.25/Mtok | 2026-09-09 |
| cost | nebius | 1.67x price · nebius/google/gemma-3-27b-it: input_price_per_mtok $0.06/Mtok to $0.10/Mtok | 2026-09-06 |
| cost | qwen | 1.67x price · qwen/qwen3.5-35b-a3b: output_price_per_mtok $0.75/Mtok to $1.25/Mtok | 2026-09-05 |
| cost | deepseek | 1.65x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.63/Mtok to $1.04/Mtok | 2026-09-07 |
| cost | deepseek | 1.65x price · deepseek/deepseek-v4-pro: output_price_per_mtok $1.25/Mtok to $2.07/Mtok | 2026-09-07 |
| cost | deepseek | 1.65x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-09-07 |
| cost | ~deepseek | 1.60x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.10/Mtok to $0.16/Mtok | 2026-09-02 |
| cost | tencent | 1.60x price · tencent/hy3: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +14 aliases | 2026-09-10 |
| cost | tencent | 1.60x price · tencent/hy3: output_price_per_mtok $0.33/Mtok to $0.53/Mtok +14 aliases | 2026-09-10 |
| cost | tencent | 1.60x price · tencent/hy3: input_price_per_mtok $0.08/Mtok to $0.13/Mtok +14 aliases | 2026-09-10 |
| cost | qwen | 1.60x price · qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.55/Mtok to $0.88/Mtok | 2026-09-09 |
| cost | openrouter | 1.60x price · openrouter/nvidia/nemotron-3.5-lightning: input_price_per_mtok $0.05/Mtok to $0.08/Mtok | 2026-09-06 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.66/Mtok to $1.05/Mtok | 2026-09-08 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.15/Mtok | 2026-09-08 |
| cost | deepseek | 1.59x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-08 |
| cost | deepseek | 1.57x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.09/Mtok to $0.14/Mtok | 2026-09-01 |
| cost | qwen | 1.57x price · qwen/qwen3-235b-a22b-2507: output_price_per_mtok $0.35/Mtok to $0.55/Mtok +1 alias | 2026-09-05 |
| cost | fireworks_ai | 1.57x price · fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_price_per_mtok $0.14/Mtok to $0.22/Mtok +1 alias | 2026-09-01 |
| cost | nvidia | 1.56x price · nvidia/nemotron-3-ultra-550b-a55b: cache_read_price_per_mtok $0.12/Mtok to $0.19/Mtok | 2026-09-04 |
| cost | qwen | 1.55x price · qwen/qwen3-30b-a3b-instruct-2507: output_price_per_mtok $0.19/Mtok to $0.30/Mtok +1 alias | 2026-09-10 |
| cost | deepseek | 1.55x price · deepseek/deepseek-v4-pro: output_price_per_mtok $2.06/Mtok to $3.20/Mtok | 2026-09-01 |
| cost | deepseek | 1.55x price · deepseek/deepseek-v4-pro: input_price_per_mtok $1.03/Mtok to $1.60/Mtok | 2026-09-01 |
| cost | openrouter | 1.50x price · openrouter/openai/gpt-oss-20b: input_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-06 |
| cost | mistral | 1.50x price · mistral/magistral-medium-latest: output_price_per_mtok $5.00/Mtok to $7.50/Mtok | 2026-09-03 |
| cost | openai | 1.50x price · gpt-realtime-2: output_price_per_mtok $16.00/Mtok to $24.00/Mtok | 2026-09-03 |
| cost | nebius | 1.50x price · nebius/google/gemma-3-27b-it: output_price_per_mtok $0.20/Mtok to $0.30/Mtok | 2026-09-06 |
| cost | qwen | 1.50x price · qwen/qwen3.5-397b-a17b: output_price_per_mtok $2.34/Mtok to $3.50/Mtok +1 alias | 2026-09-08 |
| cost | ~deepseek | 1.44x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0090/Mtok to $0.01/Mtok | 2026-09-07 |
| cost | nvidia | 1.42x price · nvidia/nemotron-3-ultra-550b-a55b: output_price_per_mtok $2.20/Mtok to $3.12/Mtok | 2026-09-02 |
| cost | qwen | 1.41x price · qwen/qwen3.5-397b-a17b: input_price_per_mtok $0.39/Mtok to $0.55/Mtok +1 alias | 2026-09-08 |
| cost | minimax | 1.38x price · minimax/minimax-m1: input_price_per_mtok $0.40/Mtok to $0.55/Mtok +1 alias | 2026-09-09 |
| cost | z-ai | 1.38x price · z-ai/glm-4.6: cache_read_price_per_mtok $0.08/Mtok to $0.11/Mtok +1 alias | 2026-09-08 |
| cost | openrouter | 1.38x price · openrouter/z-ai/glm-4.6: input_price_per_mtok $0.40/Mtok to $0.55/Mtok | 2026-09-06 |
| cost | openrouter | 1.36x price · openrouter/qwen/qwen3-coder: input_price_per_mtok $0.22/Mtok to $0.30/Mtok | 2026-09-06 |
| cost | openrouter | 1.35x price · openrouter/deepseek/deepseek-v3.2-exp: input_price_per_mtok $0.20/Mtok to $0.27/Mtok | 2026-09-06 |
| cost | qwen | 1.33x price · qwen/qwen2.5-vl-72b-instruct: output_price_per_mtok $0.75/Mtok to $1.00/Mtok +1 alias | 2026-09-03 |
| cost | nvidia | 1.30x price · nvidia/nemotron-3-ultra-550b-a55b: output_price_per_mtok $2.40/Mtok to $3.12/Mtok | 2026-09-04 |
| cost | ~deepseek | 1.30x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0100/Mtok to $0.01/Mtok | 2026-09-02 |
| cost | openrouter | 1.30x price · openrouter/openai/gpt-oss-20b: output_price_per_mtok $0.10/Mtok to $0.13/Mtok | 2026-09-06 |
| cost | qwen | 1.29x price · qwen/qwen3.6-35b-a3b: output_price_per_mtok $0.70/Mtok to $0.90/Mtok | 2026-09-03 |
| cost | z-ai | 1.28x price · z-ai/glm-4.6: input_price_per_mtok $0.43/Mtok to $0.55/Mtok +1 alias | 2026-09-08 |
| cost | openrouter | 1.27x price · openrouter/deepseek/deepseek-r1: input_price_per_mtok $0.55/Mtok to $0.70/Mtok | 2026-09-06 |
| cost | deepseek | 1.27x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.07/Mtok to $0.09/Mtok | 2026-09-10 |
| cost | deepseek | 1.27x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.14/Mtok to $0.18/Mtok | 2026-09-10 |
| cost | deepseek | 1.27x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-09-10 |
| cost | z-ai | 1.26x price · z-ai/glm-4.6: output_price_per_mtok $1.75/Mtok to $2.20/Mtok +1 alias | 2026-09-08 |
| cost | openrouter | 1.26x price · openrouter/z-ai/glm-4.6: output_price_per_mtok $1.75/Mtok to $2.20/Mtok | 2026-09-06 |
| truncation | qwen | -20% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 81,920 to 65,536 | 2026-09-10 |
| truncation | ~z-ai | -20% ctx · ~z-ai/glm-latest: context_tokens cut 1,310,720 to 1,048,576 inferred | 2026-09-09 |
| truncation | z-ai | -20% ctx · z-ai/glm-5.3: context_tokens cut 1,310,720 to 1,048,576 | 2026-09-09 |
| truncation | nebius | -14% max out · nebius/google/gemma-3-27b-it: max_output_tokens cut 128,000 to 110,000 | 2026-09-06 |
| truncation | nebius | -14% ctx · nebius/google/gemma-3-27b-it: context_tokens cut 128,000 to 110,000 | 2026-09-06 |
| truncation | z-ai | -2% max out · z-ai/glm-5.1: max_output_tokens cut 131,072 to 128,000 +7 aliases | 2026-09-10 |
| truncation | deepseek | -2% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 393,216 to 384,000 +7 aliases | 2026-09-10 |
| truncation | deepseek | -2% max out · deepseek/deepseek-chat: max_output_tokens cut 16,384 to 16,000 | 2026-09-10 |
| truncation | z-ai | -1% ctx · z-ai/glm-4.7-flash: context_tokens cut 202,752 to 200,000 | 2026-09-10 |
| capability | azure | lost max_reasoning_effort · azure/gpt-6-astra: capabilities lost max_reasoning_effort +1 alias | 2026-09-09 |
| capability | nex-agi | lost frequency_penalty · nex-agi/nex-n2-pro: capabilities lost frequency_penalty inferred | 2026-09-07 |
| capability | novita | lost vision · novita/openai/gpt-oss-120b: capabilities lost vision +1 alias | 2026-09-06 |
| capability | bedrock_converse | lost native_structured_output · anthropic.claude-fable-5-1: capabilities lost native_structured_output +15 aliases | 2026-09-02 |
| capability | ~anthropic | lost tool_choice · ~anthropic/claude-fable-latest: capabilities lost tool_choice inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-pro is no longer listed by this source — calls to it are expected to fail | 2026-09-08 |
| availability | nex-agi | nex-agi/nex-n2-mini is no longer listed by this source — calls to it are expected to fail | 2026-09-08 |
| availability | gemini-flash-latest-high-res-exp is no longer listed by this source — calls to it are expected to fail | 2026-09-04 | |
| availability | nvidia | nvidia/nemotron-3-ultra-550b-a55b:batch is no longer listed by this source — calls to it are expected to fail | 2026-09-03 |
| availability | anthropic | claude-2.0 is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-01 |
| availability | anthropic | claude-2.1 is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-01 |
| availability | anthropic | anthropic/claude-opus-4.7-fast is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | anthropic | anthropic/claude-opus-4.8-fast is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | anthropic | anthropic/claude-opus-5-fast is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | mistralai | mistralai/codestral-2508:batch is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | mistralai | mistralai/mistral-large-2512:batch is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | mistralai | mistralai/mistral-small-2603:batch is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | mistralai | mistralai/ministral-8b-2512:batch is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | mistralai | mistralai/mistral-medium-3.1:batch is no longer listed by this source — calls to it are expected to fail | 2026-09-01 |
| availability | inception | inception/mercury-2.5-preview vanished; inception/mercury-2.5 appeared with identical price and context — most likely a rename inferred | 2026-09-08 |
| availability | qwen | qwen/qwen3.8-max vanished; qwen/qwen3.8-max-0902 appeared with identical price and context — most likely a rename inferred | 2026-09-05 |
| availability | gigachat | gigachat/GigaChat-2-Lite vanished; gigachat/GigaChat-2 appeared with identical price and context — most likely a rename inferred | 2026-09-01 |
| availability | nousresearch | [one catalog only — litellm still list it] nousresearch/hermes-4-70b is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-09 |
| availability | minimax | [one catalog only — litellm still list it] minimax/minimax-m3:free is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-07 |
| availability | minimax | [one catalog only — litellm still list it] minimax/minimax-m2.7:free is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-07 |
| availability | z-ai | [one catalog only — litellm still list it] z-ai/glm-5.2:free is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-06 |
| cost | openrouter | openrouter/xiaomi/mimo-v2-flash: cache_read_price_per_mtok $0.0000/Mtok to $0.01/Mtok +1 alias | 2026-09-06 |
| cost | openrouter | openrouter/minimax/minimax-m2.1: cache_read_price_per_mtok $0.0000/Mtok to $0.03/Mtok | 2026-09-06 |
| cost | openrouter | openrouter/z-ai/glm-4.7: cache_read_price_per_mtok $0.0000/Mtok to $0.08/Mtok | 2026-09-06 |
| availability | ibm-granite | [one catalog only — litellm still list it] ibm-granite/granite-4.1-8b is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-04 |
| availability | azure_ai | [one catalog only — litellm, openrouter still list it] azure_ai/deepseek-v4-flash-0731 is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-03 |
| availability | anthropic | [one catalog only — litellm still list it] claude-3-sonnet-20240229 is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-01 |
| availability | [one catalog only — litellm still list it] gemini-robotics-er-1.6-preview is no longer listed by this source — calls to it are expected to fail inferred | 2026-09-01 |
Announced this period
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| availability | scaleway | scaleway/mistralai/pixtral-12b-2409: retires 2026-10-01 (in 21 days) +1 alias | 2026-09-10 |
| availability | scaleway | scaleway/mistralai/pixtral-12b-2409: deprecation announced for 2026-10-01 — in 21 days +1 alias | 2026-09-10 |
| availability | bedrock_converse | us-gov.anthropic.claude-3-haiku-20240307-v1:0: retires 2026-09-10 (in 1 days) inferred | 2026-09-09 |
| availability | azure | azure/gpt-realtime-2: retires 2026-08-31 (6 days ago) inferred | 2026-09-06 |
| availability | databricks | databricks/databricks-gemini-2-5-flash: deprecation announced for 2026-10-02 — in 26 days inferred | 2026-09-06 |
| availability | databricks | databricks/databricks-gemini-2-5-flash: retires 2026-10-02 (in 26 days) inferred | 2026-09-06 |
| availability | z-ai | z-ai/glm-4.7-flash: retirement announced for 2026-09-10 — in 6 days | 2026-09-04 |
| availability | z-ai | z-ai/glm-4.7-flash: retires 2026-09-10 (in 6 days) | 2026-09-04 |
| availability | azure_ai | azure_ai/MAI-Image-2.5-Flash: retires 2026-10-01 (in 28 days) +1 alias inferred | 2026-09-03 |
| availability | azure_ai | azure_ai/MAI-Image-2.5-Flash: deprecation announced for 2026-10-01 — in 28 days +1 alias inferred | 2026-09-03 |
| availability | cerebras | cerebras/zai-glm-4.7: retires 2026-08-17 (15 days ago) inferred | 2026-09-01 |
| availability | bedrock | amazon.nova-sonic-v1:0: retires 2026-09-14 (in 13 days) inferred | 2026-09-01 |
| availability | cerebras | cerebras/zai-glm-4.7: deprecation announced for 2026-08-17 (date already passed) inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-mini: retires 2026-09-08 (in 7 days) +1 alias inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-mini: retirement announced for 2026-09-08 — in 7 days +1 alias inferred | 2026-09-01 |
| availability | azure_ai | azure_ai/gpt-chat-latest: retires 2026-12-02 (in 83 days) | 2026-09-10 |
| availability | azure_ai | azure_ai/codex-mini: retires 2026-11-15 (in 66 days) inferred | 2026-09-10 |
| availability | databricks | databricks/databricks-claude-sonnet-4: retires 2026-10-09 (in 33 days) inferred | 2026-09-06 |
| availability | databricks | databricks/databricks-claude-sonnet-4: deprecation announced for 2026-10-09 — in 33 days inferred | 2026-09-06 |
| availability | azure_ai | azure_ai/kimi-k2.7-code: retires 2026-10-03 (in 30 days) | 2026-09-03 |
| availability | azure_ai | azure_ai/claude-opus-4-7: retires 2027-04-06 (in 228 days) +2 aliases | 2026-09-10 |
| availability | azure_ai | azure_ai/model_router: deprecation announced for 2027-05-20 — in 252 days | 2026-09-10 |
| availability | azure | azure/eu/gpt-5.5-2026-04-23: retires 2027-10-26 (in 411 days) +5 aliases | 2026-09-10 |
| availability | azure | azure/eu/gpt-5.5-2026-04-23: deprecation announced for 2027-10-26 — in 411 days +5 aliases | 2026-09-10 |
| availability | azure_ai | azure_ai/model-router: retires 2027-05-20 (in 252 days) +1 alias | 2026-09-10 |
| availability | azure_ai | azure_ai/Cohere-parse-v5: retires 2026-12-15 (in 100 days) +1 alias | 2026-09-10 |
| availability | z-ai | z-ai/glm-5.3-flash: retirement announced for 2098-12-31 — in 26410 days | 2026-09-10 |
| availability | azure_ai | azure_ai/claude-fable-5-1: retires 2027-12-05 (in 455 days) +1 alias | 2026-09-06 |
| availability | azure | azure/gpt-4.1-nano-2025-04-14: deprecation announced for 2027-04-14 — in 229 days +2 aliases | 2026-09-06 |
| availability | anthropic | claude-fable-5-1: retires 2027-09-01 (in 365 days) +1 alias | 2026-09-06 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-fable-5-1: retires 2027-03-01 (in 176 days) +1 alias | 2026-09-06 |
| availability | azure_ai | azure_ai/claude-fable-5-1: deprecation announced for 2027-12-05 — in 455 days +1 alias | 2026-09-06 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-fable-5-1: deprecation announced for 2027-03-01 — in 176 days +1 alias | 2026-09-06 |
| availability | azure | azure/eu/gpt-4o-2024-08-06: retires 2027-04-14 (in 245 days) +23 aliases | 2026-09-06 |
| availability | azure | azure/gpt-realtime-2.1-mini: retires 2027-06-25 (in 292 days) +1 alias | 2026-09-06 |
| availability | azure_ai | azure_ai/DeepSeek-V4-Flash-0731: retires 2026-12-03 (in 91 days) +1 alias | 2026-09-03 |
| availability | z-ai | z-ai/glm-4.5v: retirement announced for 2026-12-31 — in 120 days +3 aliases | 2026-09-02 |
| availability | ~z-ai | ~z-ai/glm-flash-latest: retires 2098-12-31 (in 26418 days) +1 alias | 2026-09-02 |
| availability | openai | gpt-4o-mini-transcribe: deprecation announced for 2027-02-26 — in 178 days +3 aliases | 2026-09-01 |
| availability | azure_ai | azure_ai/claude-opus-5: deprecation announced for 2027-07-08 — in 310 days | 2026-09-01 |
| availability | azure_ai | azure_ai/claude-opus-4-8: retires 2027-09-01 (in 365 days) | 2026-09-01 |
| availability | bedrock_converse | anthropic.claude-opus-4-1-20250805-v1:0: retires 2027-01-08 (in 149 days) +2 aliases | 2026-09-01 |
| availability | azure_ai | azure_ai/claude-sonnet-5: deprecation announced for 2027-06-30 — in 302 days | 2026-09-01 |
| availability | bedrock_converse | anthropic.claude-opus-4-1-20250805-v1:0: deprecation announced for 2027-01-08 — in 149 days +2 aliases | 2026-09-01 |
| availability | vertex_ai-language-models | gemini-live-2.5-flash-native-audio: retires 2026-12-13 (in 103 days) | 2026-09-01 |
| availability | vertex_ai-embedding-models | text-embedding-004: deprecation announced for 2027-04-01 — in 212 days | 2026-09-01 |
| availability | azure_ai | azure_ai/claude-opus-5: retires 2027-07-08 (in 310 days) | 2026-09-01 |
| availability | vertex_ai-language-models | gemini-live-2.5-flash-native-audio: deprecation announced for 2026-12-13 — in 103 days | 2026-09-01 |
| availability | azure_ai | azure_ai/claude-sonnet-5: retires 2027-06-30 (in 302 days) | 2026-09-01 |
| availability | vertex_ai-embedding-models | multimodalembedding@001: retires 2027-04-01 (in 223 days) +3 aliases | 2026-09-01 |
| availability | azure_ai | azure_ai/claude-opus-4-8: deprecation announced for 2027-09-01 — in 365 days | 2026-09-01 |
Who changed the most
| Provider | Unannounced changes |
|---|---|
| deepseek | 35 |
| openrouter | 29 |
| qwen | 24 |
| z-ai | 17 |
| nvidia | 8 |
| nebius | 8 |
| mistralai | 7 |
| ~z-ai | 6 |
| ~deepseek | 6 |
| anthropic | 6 |
| 5 | |
| watsonx | 4 |