400
changes shipped in the last two months that nobody announced. Prices moved, context windows shrank, capabilities disappeared — no email, no changelog entry, no error at runtime.
345
price changes
you pay more for the same call
130
capacity cuts
prompts truncate, no error
39
capabilities removed
a feature quietly went away
203
announced retirements
these ones you were told about
Evidence behind the numbers: 66 models confirmed against the provider's own API, 123 with dates taken from the provider's own documentation (extracted by a model, so labelled inferred), the remaining 2287 from third-party catalogs only. Every model page says which of the three it rests on.
Biggest silent changes
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| truncation | thedrummer | -97% max out · thedrummer/unslopnemo-12b: max_output_tokens cut 1,024,000 to 32,768 inferred | 2026-08-24 |
| truncation | openai | -97% max out · openai/gpt-4.1-nano: max_output_tokens cut 942,818 to 32,768 | 2026-08-30 |
| truncation | nvidia | -96% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 461,059 to 16,384 | 2026-08-28 |
| truncation | qwen | -94% max out · qwen/qwen3-next-80b-a3b-instruct: max_output_tokens cut 262,144 to 16,384 +1 alias | 2026-08-19 |
| truncation | watsonx | -94% max out · watsonx/meta-llama/llama-4-maverick-17b: max_output_tokens cut 128,000 to 8,192 inferred | 2026-09-06 |
| truncation | openrouter | -94% max out · openrouter/anthropic/claude-sonnet-4.5: max_output_tokens cut 1,000,000 to 64,000 | 2026-09-06 |
| truncation | vertex_ai | -94% ctx · vertex_ai/xai/grok-4.1-fast-non-reasoning: context_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| truncation | vertex_ai | -94% max out · vertex_ai/xai/grok-4.1-fast-non-reasoning: max_output_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| cost | deepseek | 15.00x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.0036/Mtok to $0.05/Mtok | 2026-08-11 |
| truncation | thinkingmachines | -93% max out · thinkingmachines/inkling: max_output_tokens cut 471,859 to 32,768 | 2026-09-10 |
| truncation | qwen | -93% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 235,929 to 16,384 +5 aliases | 2026-09-09 |
| cost | deepseek | 12.14x price · deepseek-v4-pro: cache_read_price_per_mtok $0.0036/Mtok to $0.04/Mtok +1 alias | 2026-08-21 |
| cost | z-ai | 11.00x price · z-ai/glm-5.2: output_price_per_mtok $0.22/Mtok to $2.42/Mtok | 2026-08-10 |
| cost | z-ai | 10.86x price · z-ai/glm-5.2: input_price_per_mtok $0.07/Mtok to $0.76/Mtok | 2026-08-10 |
| cost | z-ai | 10.77x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.01/Mtok to $0.14/Mtok | 2026-08-10 |
| truncation | novita | -91% ctx · novita/meta-llama/llama-3.3-70b-instruct: context_tokens cut 131,072 to 12,288 | 2026-08-28 |
| cost | xai | 10.00x price · xai/grok-code-fast-1-0825: cache_read_price_per_mtok $0.02/Mtok to $0.20/Mtok +2 aliases | 2026-08-12 |
| capability | arcee-ai | lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs · arcee-ai/trinity-large-thinking: capabilities lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs inferred | 2026-08-31 |
All 400 changes → · Check your own models →
Announced retirements
The part providers do tell you about. Included for completeness — it is the smaller half of the problem.
| Model | Vendor | Sunset | Remaining | Severity |
|---|---|---|---|---|
| gov-east-1/anthropic.claude-3-haiku-20240307 | gov-east-1 | T-0d | critical | |
| gov-west-1/anthropic.claude-3-haiku-20240307 | gov-west-1 | T-0d | critical | |
| amazon/gov.anthropic.claude-3-haiku-20240307 | amazon | T-0d | critical | |
| amazon/nova-sonic | amazon | T-4d | critical | |
| amazon/nova-premier | amazon | T-4d | critical | |
| openai/sora-2 | openai | T-14d | critical | |
| openai/sora-2-2025-10-06 | openai | T-14d | critical | |
| openai/sora-2-2025-12-08 | openai | T-14d | critical | |
| openai/sora-2-pro | openai | T-14d | critical | |
| openai/sora-2-pro-2025-10-06 | openai | T-14d | critical |