Unannounced changes to AI APIs
Price moves, context-window cuts, removed capabilities and vanished models — grouped so one upstream change is one row, and ranked by how big the change is rather than when we spotted it.
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| truncation | gemini | -99% ctx · gemini/gemini-2.5-pro-preview-tts: context_tokens cut 1,048,576 to 8,192 | 2026-09-16 |
| truncation | openrouter | -98% max out · openrouter/moonshotai/kimi-k3:batch: max_output_tokens cut 943,718 to 16,384 | 2026-09-22 |
| truncation | bedrock_converse | -97% max out · zai.glm-4.7-flash: max_output_tokens cut 128,000 to 4,000 +1 alias | 2026-09-21 |
| cost | deepseek | 31.35x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0010/Mtok to $0.03/Mtok | 2026-09-27 |
| cost | ~deepseek | 31.35x price · ~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.0010/Mtok to $0.03/Mtok | 2026-09-27 |
| cost | openrouter | 31.35x price · openrouter/deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0010/Mtok to $0.03/Mtok | 2026-09-28 |
| truncation | thedrummer | -97% max out · thedrummer/unslopnemo-12b: max_output_tokens cut 1,024,000 to 32,768 inferred | 2026-08-24 |
| truncation | openai | -97% max out · openai/gpt-4.1-nano: max_output_tokens cut 942,818 to 32,768 | 2026-08-30 |
| truncation | nvidia | -96% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 461,059 to 16,384 | 2026-08-28 |
| cost | openrouter | 28.12x price · openrouter/deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.0088/Mtok to $0.25/Mtok | 2026-09-27 |
| cost | deepseek | 27.79x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.0088/Mtok to $0.24/Mtok | 2026-09-27 |
| truncation | azure_ai | -96% max out · azure_ai/grok-4.3: max_output_tokens cut 200,000 to 8,192 | 2026-09-28 |
| cost | ~deepseek | 22.14x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.0078/Mtok to $0.17/Mtok | 2026-09-27 |
| truncation | bedrock_converse | -95% max out · deepseek.v3.2: max_output_tokens cut 163,840 to 8,000 inferred | 2026-09-21 |
| cost | deepseek | 20.00x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0030/Mtok to $0.06/Mtok | 2026-09-25 |
| cost | ~deepseek | 20.00x price · ~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.0010/Mtok to $0.02/Mtok | 2026-09-29 |
| cost | ~deepseek | 19.64x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.01/Mtok to $0.25/Mtok | 2026-09-23 |
| cost | qwen | 16.87x price · qwen/qwen3.8-27b: input_price_per_mtok $0.02/Mtok to $0.42/Mtok | 2026-09-30 |
| truncation | bedrock_converse | -94% max out · moonshotai.kimi-k2.5: max_output_tokens cut 262,144 to 16,000 | 2026-09-21 |
| truncation | qwen | -94% max out · qwen/qwen3-next-80b-a3b-instruct: max_output_tokens cut 262,144 to 16,384 +1 alias | 2026-08-19 |
| truncation | azure_ai | -94% max out · azure_ai/grok-4: max_output_tokens cut 131,072 to 8,192 +1 alias inferred | 2026-09-28 |
| truncation | watsonx | -94% max out · watsonx/meta-llama/llama-4-maverick-17b: max_output_tokens cut 128,000 to 8,192 inferred | 2026-09-06 |
| truncation | openrouter | -94% max out · openrouter/anthropic/claude-sonnet-4.5: max_output_tokens cut 1,000,000 to 64,000 | 2026-09-06 |
| truncation | vertex_ai | -94% ctx · vertex_ai/xai/grok-4.1-fast-non-reasoning: context_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| truncation | vertex_ai | -94% max out · vertex_ai/xai/grok-4.1-fast-non-reasoning: max_output_tokens cut 2,000,000 to 128,000 +1 alias inferred | 2026-09-10 |
| cost | deepseek | 15.00x price · deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.02/Mtok to $0.30/Mtok | 2026-09-28 |
| truncation | thinkingmachines | -93% max out · thinkingmachines/inkling: max_output_tokens cut 471,859 to 32,768 | 2026-09-10 |
| truncation | qwen | -93% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 235,929 to 16,384 +5 aliases | 2026-09-09 |
| truncation | openrouter | -93% max out · openrouter/qwen/qwen3-next-80b-a3b-instruct: max_output_tokens cut 235,929 to 16,384 | 2026-09-18 |
| truncation | novita | -91% ctx · novita/meta-llama/llama-3.3-70b-instruct: context_tokens cut 131,072 to 12,288 | 2026-08-28 |
| capability | arcee-ai | lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs · arcee-ai/trinity-large-thinking: capabilities lost frequency_penalty, logit_bias, logprobs, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, top_logprobs inferred | 2026-08-31 |
| cost | deepseek | 10.00x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0010/Mtok to $0.01/Mtok | 2026-09-28 |
| cost | ~deepseek | 10.00x price · ~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.0010/Mtok to $0.01/Mtok | 2026-09-28 |
| truncation | novita | -90% max out · novita/meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 120,000 to 12,288 | 2026-08-28 |
| truncation | ~z-ai | -89% max out · ~z-ai/glm-flash-latest: max_output_tokens cut 943,718 to 102,400 | 2026-09-21 |
| cost | openrouter | 8.61x price · openrouter/deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.03/Mtok to $0.30/Mtok | 2026-09-29 |
| truncation | together_ai | -88% max out · together_ai/zai-org/GLM-5.2: max_output_tokens cut 1,048,575 to 128,000 +1 alias | 2026-08-30 |
| truncation | mistralai | -88% max out · mistralai/mistral-medium-3-5:batch: max_output_tokens cut 209,715 to 26,214 inferred | 2026-09-04 |
| truncation | fireworks_ai | -88% max out · fireworks_ai/accounts/fireworks/models/kimi-k2p5: max_output_tokens cut 262,144 to 32,768 +9 aliases inferred | 2026-07-31 |
| truncation | qwen | -88% max out · qwen/qwen3-coder-30b-a3b-instruct: max_output_tokens cut 262,144 to 32,768 +2 aliases | 2026-08-13 |
| capability | lost frequency_penalty, logprobs, presence_penalty, repetition_penalty, stop, structured_outputs, top_k, top_logprobs · google/gemma-4-26b-a4b-it:free: capabilities lost frequency_penalty, logprobs, presence_penalty, repetition_penalty, stop, structured_outputs, top_k, top_logprobs inferred | 2026-08-20 | |
| truncation | mistralai | -88% ctx · mistralai/mistral-medium-3-5:batch: context_tokens cut 262,144 to 32,768 inferred | 2026-09-04 |
| truncation | vertex_ai-language-models | -88% ctx · gemini-live-2.5-flash-native-audio: context_tokens cut 1,048,576 to 131,072 inferred | 2026-09-10 |
| truncation | z-ai | -88% max out · z-ai/glm-4.6: max_output_tokens cut 131,072 to 16,384 +3 aliases | 2026-09-19 |
| truncation | aion-labs | -88% ctx · aion-labs/aion-2.0: context_tokens cut 1,048,576 to 131,072 +2 aliases | 2026-09-22 |
| truncation | openrouter | -88% ctx · openrouter/aion-labs/aion-2.0: context_tokens cut 1,048,576 to 131,072 +2 aliases | 2026-09-23 |
| truncation | gemini | -88% ctx · gemini-2.5-flash-native-audio-latest: context_tokens cut 1,048,576 to 131,072 +4 aliases | 2026-09-26 |
| truncation | azure | -88% ctx · azure/eu/gpt-4o-realtime-preview-2024-12-17: context_tokens cut 128,000 to 16,000 +2 aliases | 2026-09-30 |
| capability | deepseek | lost logit_bias, logprobs, min_p, presence_penalty, repetition_penalty, seed, stop, top_logprobs · deepseek/deepseek-v3.2-exp: capabilities lost logit_bias, logprobs, min_p, presence_penalty, repetition_penalty, seed, stop, top_logprobs | 2026-09-30 |
| truncation | openrouter | -87% max out · openrouter/z-ai/glm-4.6: max_output_tokens cut 131,000 to 16,384 | 2026-09-18 |
| cost | ~deepseek | 7.94x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.01/Mtok to $0.09/Mtok | 2026-09-25 |
| truncation | openrouter | -87% max out · openrouter/z-ai/glm-5.2:free: max_output_tokens cut 230,400 to 29,491 | 2026-09-18 |
| truncation | meta-llama | -87% max out · meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 128,000 to 16,384 | 2026-08-04 |
| truncation | openrouter | -87% ctx · openrouter/z-ai/glm-5.2:free: context_tokens cut 256,000 to 32,768 | 2026-09-18 |
| truncation | openrouter | -87% max out · openrouter/mistralai/mistral-small-3.2-24b-instruct: max_output_tokens cut 128,000 to 16,384 | 2026-09-18 |
| truncation | ~z-ai | -86% max out · ~z-ai/glm-latest: max_output_tokens cut 943,718 to 128,000 +1 alias inferred | 2026-09-11 |
| truncation | qwen | -86% max out · qwen/qwen3-30b-a3b-instruct-2507: max_output_tokens cut 235,929 to 32,000 +3 aliases | 2026-09-29 |
| cost | z-ai | 7.37x price · z-ai/glm-5.3: input_price_per_mtok $0.19/Mtok to $1.40/Mtok | 2026-09-29 |
| truncation | ~deepseek | -86% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 943,718 to 131,072 +1 alias inferred | 2026-09-13 |
| truncation | z-ai | -86% max out · z-ai/glm-5.3-flash: max_output_tokens cut 943,718 to 131,072 +3 aliases | 2026-09-24 |
| truncation | deepseek | -86% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 943,718 to 131,072 +1 alias | 2026-09-26 |
| truncation | openrouter | -86% max out · openrouter/z-ai/glm-5.3-flash: max_output_tokens cut 943,718 to 131,072 +3 aliases | 2026-09-27 |
| truncation | ~z-ai | -86% max out · ~z-ai/glm-flash-latest: max_output_tokens cut 943,718 to 131,072 +7 aliases | 2026-09-28 |
| truncation | openrouter | -86% max out · openrouter/z-ai/glm-5.3: max_output_tokens cut 943,717 to 131,072 | 2026-09-19 |
| truncation | z-ai | -86% max out · z-ai/glm-5.3: max_output_tokens cut 943,717 to 131,072 +3 aliases | 2026-09-29 |
| truncation | qwen | -86% max out · qwen/qwen3-next-80b-a3b-thinking: max_output_tokens cut 235,929 to 32,768 +1 alias | 2026-09-20 |
| truncation | openrouter | -86% max out · openrouter/nvidia/nemotron-3.5-lightning: max_output_tokens cut 235,929 to 32,768 +1 alias | 2026-09-29 |
| truncation | meta | -86% max out · meta/muse-glimmer-30b: max_output_tokens cut 117,964 to 16,384 +1 alias | 2026-09-20 |
| truncation | openrouter | -86% max out · openrouter/meta/muse-glimmer-30b: max_output_tokens cut 117,964 to 16,384 | 2026-09-20 |
| truncation | meta-llama | -86% max out · meta-llama/llama-3.3-70b-instruct: max_output_tokens cut 115,200 to 16,384 +1 alias | 2026-09-14 |
| truncation | openrouter | -86% max out · openrouter/meta-llama/llama-4-maverick: max_output_tokens cut 115,200 to 16,384 | 2026-09-18 |
| truncation | meta-llama | -84% ctx · meta-llama/llama-guard-4-12b: context_tokens cut 1,048,576 to 163,840 | 2026-08-26 |
| capability | tencent | lost frequency_penalty, presence_penalty, repetition_penalty, response_format, stop, structured_outputs · tencent/hy3-preview: capabilities lost frequency_penalty, presence_penalty, repetition_penalty, response_format, stop, structured_outputs inferred | 2026-08-03 |
| cost | ~deepseek | 6.00x price · ~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.06/Mtok +1 alias | 2026-09-25 |
| truncation | deepseek | -83% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 384,000 to 65,536 inferred | 2026-08-01 |
| cost | z-ai | 5.78x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.04/Mtok to $0.26/Mtok | 2026-09-27 |
| truncation | nvidia | -82% max out · nvidia/nemotron-3-ultra-550b-a55b: max_output_tokens cut 182,520 to 32,768 +2 aliases | 2026-09-17 |
| cost | z-ai | 5.12x price · z-ai/glm-5.3: input_price_per_mtok $0.27/Mtok to $1.40/Mtok | 2026-09-27 |
| truncation | bedrock_converse | -80% ctx · minimax.minimax-m2.5: context_tokens cut 1,000,000 to 196,000 inferred | 2026-09-22 |
| truncation | deepseek | -80% max out · deepseek/deepseek-v3.1-terminus: max_output_tokens cut 163,840 to 32,768 | 2026-08-20 |
| truncation | openrouter | -80% max out · openrouter/deepseek/deepseek-chat-v3.1: max_output_tokens cut 163,840 to 32,768 | 2026-09-18 |
| truncation | anthropic | -80% ctx · anthropic/claude-sonnet-4: context_tokens cut 1,000,000 to 200,000 | 2026-09-21 |
| truncation | openrouter | -80% ctx · openrouter/anthropic/claude-sonnet-4.5: context_tokens cut 1,000,000 to 200,000 +1 alias | 2026-09-21 |
| capability | deepseek | lost logit_bias, min_p, presence_penalty, repetition_penalty, stop · deepseek/deepseek-chat-v3-0324: capabilities lost logit_bias, min_p, presence_penalty, repetition_penalty, stop | 2026-09-29 |
| cost | ~deepseek | 4.78x price · ~deepseek/deepseek-pro-latest: output_price_per_mtok $0.73/Mtok to $3.50/Mtok | 2026-09-27 |
| cost | ~z-ai | 4.59x price · ~z-ai/glm-latest: output_price_per_mtok $0.56/Mtok to $2.57/Mtok | 2026-09-27 |
| cost | z-ai | 4.59x price · z-ai/glm-5.3: output_price_per_mtok $0.56/Mtok to $2.57/Mtok | 2026-09-27 |
| truncation | deepseek | -77% max out · deepseek/deepseek-chat-v3.1: max_output_tokens cut 144,900 to 32,768 +1 alias | 2026-09-08 |
| cost | openrouter | 4.42x price · openrouter/deepseek/deepseek-v4-pro-0813: output_price_per_mtok $0.79/Mtok to $3.50/Mtok | 2026-09-27 |
| cost | deepseek | 4.42x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $0.79/Mtok to $3.50/Mtok | 2026-09-27 |
| cost | qwen | 4.27x price · qwen/qwen3.8-27b: cache_read_price_per_mtok $0.02/Mtok to $0.09/Mtok | 2026-09-30 |
| truncation | nebius | -76% ctx · nebius/Qwen/Qwen2.5-VL-72B-Instruct: context_tokens cut 131,072 to 32,000 | 2026-09-06 |
| truncation | nebius | -76% max out · nebius/Qwen/Qwen2.5-VL-72B-Instruct: max_output_tokens cut 131,072 to 32,000 | 2026-09-06 |
| truncation | z-ai | -75% max out · z-ai/glm-5.3: max_output_tokens cut 943,718 to 235,929 | 2026-09-27 |
| truncation | ~z-ai | -75% max out · ~z-ai/glm-latest: max_output_tokens cut 943,718 to 235,929 +5 aliases | 2026-09-27 |
| truncation | fireworks_ai | -75% max out · fireworks_ai/accounts/fireworks/models/gpt-oss-120b: max_output_tokens cut 131,072 to 32,768 +1 alias | 2026-06-18 |
| truncation | liquid | -75% max out · liquid/lfm-2.5-2.6b:free: max_output_tokens cut 32,768 to 8,192 inferred | 2026-08-14 |
| truncation | ~deepseek | -75% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 1,048,576 to 262,144 +1 alias inferred | 2026-08-23 |
| truncation | qwen | -75% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 262,144 to 65,536 +8 aliases | 2026-08-23 |
| truncation | qwen | -75% max out · qwen/qwen2.5-vl-72b-instruct: max_output_tokens cut 115,200 to 28,800 | 2026-08-26 |
| truncation | openai | -75% ctx · gpt-realtime-mini: context_tokens cut 128,000 to 32,000 | 2026-09-03 |
| capability | vertex_ai-language-models | lost pdf_input, prompt_caching, response_schema, url_context · gemini-live-2.5-flash-native-audio: capabilities lost pdf_input, prompt_caching, response_schema, url_context inferred | 2026-09-10 |
| capability | gemini | lost function_calling, response_schema, vision, web_search · gemini/gemini-2.5-pro-preview-tts: capabilities lost function_calling, response_schema, vision, web_search | 2026-09-16 |
| capability | replicate | lost function_calling, parallel_function_calling, response_schema, tool_choice · replicate/google/gemini-2.5-flash: capabilities lost function_calling, parallel_function_calling, response_schema, tool_choice; gained reasoning, video_input | 2026-09-19 |
| capability | replicate | lost function_calling, parallel_function_calling, response_schema, tool_choice · replicate/google/gemini-3-pro: capabilities lost function_calling, parallel_function_calling, response_schema, tool_choice; gained audio_input, video_input inferred | 2026-09-19 |
| truncation | openrouter | -75% max out · openrouter/qwen/qwen3.5-35b-a3b: max_output_tokens cut 65,536 to 16,384 | 2026-09-20 |
| truncation | qwen | -75% max out · qwen/qwen3.5-35b-a3b: max_output_tokens cut 65,536 to 16,384 | 2026-09-20 |
| truncation | nvidia | -75% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 131,072 to 32,768 | 2026-09-28 |
| cost | deepseek | 4.00x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.32/Mtok to $1.28/Mtok | 2026-09-30 |
| truncation | gemini | -75% max out · gemini/gemini-2.5-pro-preview-tts: max_output_tokens cut 65,535 to 16,384 | 2026-09-16 |
| truncation | openrouter | -75% max out · openrouter/qwen/qwen3-coder: max_output_tokens cut 262,100 to 65,536 | 2026-09-18 |
| cost | z-ai | 3.94x price · z-ai/glm-5.3: output_price_per_mtok $1.12/Mtok to $4.40/Mtok | 2026-09-28 |
| cost | z-ai | 3.94x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.07/Mtok to $0.26/Mtok | 2026-09-28 |
| cost | openrouter | 3.94x price · openrouter/z-ai/glm-5.3: cache_read_price_per_mtok $0.07/Mtok to $0.26/Mtok | 2026-09-29 |
| cost | z-ai | 3.94x price · z-ai/glm-5.3: input_price_per_mtok $0.36/Mtok to $1.40/Mtok | 2026-09-28 |
| cost | openrouter | 3.94x price · openrouter/z-ai/glm-5.3: input_price_per_mtok $0.36/Mtok to $1.40/Mtok | 2026-09-29 |
| truncation | novita | -74% max out · novita/qwen/qwen2.5-7b-instruct: max_output_tokens cut 32,000 to 8,192 inferred | 2026-08-28 |
| truncation | azure | -74% ctx · azure/gpt-5.4-mini-2026-03-17: context_tokens cut 1,050,000 to 272,000 +3 aliases | 2026-07-30 |
| truncation | openai | -74% ctx · gpt-5.4-mini-2026-03-17: context_tokens cut 1,050,000 to 272,000 +3 aliases | 2026-07-30 |
| truncation | openrouter | -74% ctx · openrouter/nvidia/nemotron-3-super-120b-a12b: context_tokens cut 1,000,000 to 262,144 | 2026-09-18 |
| truncation | nvidia | -74% ctx · nvidia/nemotron-3-super-120b-a12b: context_tokens cut 1,000,000 to 262,144 +2 aliases | 2026-09-29 |
| cost | openrouter | 3.79x price · openrouter/qwen/qwen3-14b: output_price_per_mtok $0.24/Mtok to $0.91/Mtok +1 alias | 2026-09-28 |
| cost | openrouter | 3.75x price · openrouter/z-ai/glm-5.3-flash: input_price_per_mtok $0.04/Mtok to $0.15/Mtok | 2026-09-28 |
| truncation | deepseek | -72% max out · deepseek/deepseek-v4-flash-vision-exp: max_output_tokens cut 943,718 to 262,144 +1 alias | 2026-09-27 |
| truncation | openrouter | -72% max out · openrouter/deepseek/deepseek-v4-flash-vision-exp: max_output_tokens cut 943,718 to 262,144 +1 alias | 2026-09-29 |
| truncation | deepseek | -72% max out · deepseek/deepseek-v4-flash-vision-exp: max_output_tokens cut 943,717 to 262,144 | 2026-09-28 |
| truncation | qwen | -72% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 235,929 to 65,536 +4 aliases | 2026-09-29 |
| truncation | openai | -72% max out · openai/gpt-oss-20b: max_output_tokens cut 117,964 to 32,768 | 2026-09-22 |
| truncation | openrouter | -72% max out · openrouter/openai/gpt-oss-20b: max_output_tokens cut 117,964 to 32,768 | 2026-09-22 |
| cost | ~deepseek | 3.58x price · ~deepseek/deepseek-pro-latest: output_price_per_mtok $1.20/Mtok to $4.30/Mtok | 2026-09-23 |
| cost | openrouter | 3.57x price · openrouter/z-ai/glm-5.3-flash: output_price_per_mtok $0.14/Mtok to $0.50/Mtok | 2026-09-27 |
| cost | z-ai | 3.57x price · z-ai/glm-5.3-flash: output_price_per_mtok $0.14/Mtok to $0.50/Mtok +1 alias | 2026-09-28 |
| cost | ~z-ai | 3.57x price · ~z-ai/glm-flash-latest: output_price_per_mtok $0.14/Mtok to $0.50/Mtok +1 alias | 2026-09-29 |
| cost | openrouter | 3.46x price · openrouter/~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.01/Mtok to $0.04/Mtok | 2026-09-23 |
| cost | openrouter | 3.44x price · openrouter/z-ai/glm-5.3: output_price_per_mtok $0.75/Mtok to $2.57/Mtok | 2026-09-28 |
| cost | moonshotai | 3.39x price · moonshotai/kimi-k3: input_price_per_mtok $0.88/Mtok to $3.00/Mtok | 2026-09-25 |
| cost | deepseek | 3.35x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.39/Mtok to $1.32/Mtok | 2026-09-30 |
| cost | deepseek | 3.33x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0030/Mtok to $0.01/Mtok | 2026-09-23 |
| cost | openrouter | 3.33x price · openrouter/ibm-granite/granite-4.2-8b: cache_read_price_per_mtok $0.01/Mtok to $0.05/Mtok +1 alias | 2026-09-23 |
| cost | z-ai | 3.33x price · z-ai/glm-5.3-flash: input_price_per_mtok $0.04/Mtok to $0.15/Mtok | 2026-09-28 |
| cost | openrouter | 3.30x price · openrouter/~deepseek/deepseek-pro-latest: input_price_per_mtok $0.40/Mtok to $1.32/Mtok | 2026-09-23 |
| cost | openrouter | 3.30x price · openrouter/~deepseek/deepseek-pro-latest: output_price_per_mtok $1.20/Mtok to $3.96/Mtok | 2026-09-23 |
| truncation | qwen | -69% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 262,144 to 81,920 +1 alias | 2026-08-24 |
| truncation | qwen | -69% max out · qwen/qwen3.6-27b: max_output_tokens cut 262,140 to 81,920 | 2026-09-26 |
| truncation | openrouter | -69% max out · openrouter/qwen/qwen3.6-27b: max_output_tokens cut 262,140 to 81,920 | 2026-09-27 |
| truncation | openrouter | -68% max out · openrouter/anthropic/claude-haiku-4.5: max_output_tokens cut 200,000 to 64,000 | 2026-09-06 |
| capability | gemini | lost code_execution, file_search, service_tier · gemini/gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search, service_tier +1 alias | 2026-06-27 |
| capability | vertex_ai-language-models | lost code_execution, file_search, service_tier · gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search, service_tier +2 aliases | 2026-06-27 |
| capability | anthropic | lost max_completion_tokens, response_format, structured_outputs · anthropic/claude-opus-4.1: capabilities lost max_completion_tokens, response_format, structured_outputs | 2026-08-06 |
| capability | x-ai | lost frequency_penalty, presence_penalty, stop · x-ai/grok-4.3: capabilities lost frequency_penalty, presence_penalty, stop +3 aliases | 2026-08-18 |
| capability | ~x-ai | lost frequency_penalty, presence_penalty, stop · ~x-ai/grok-latest: capabilities lost frequency_penalty, presence_penalty, stop inferred | 2026-08-18 |
| capability | nvidia | lost logprobs, structured_outputs, top_logprobs · nvidia/nemotron-3-super-120b-a12b: capabilities lost logprobs, structured_outputs, top_logprobs | 2026-09-09 |
| capability | qwen | lost logit_bias, min_p, repetition_penalty · qwen/qwen3.8-flash: capabilities lost logit_bias, min_p, repetition_penalty | 2026-09-15 |
| capability | thedrummer | lost response_format, tool_choice, tools · thedrummer/unslopnemo-12b: capabilities lost response_format, tool_choice, tools inferred | 2026-09-16 |
| capability | gemini | lost function_calling, response_schema, web_search · gemini/gemini-2.5-flash-image: capabilities lost function_calling, response_schema, web_search | 2026-09-16 |
| capability | openrouter | lost function_calling, response_schema, tool_choice · openrouter/z-ai/glm-5.2:free: capabilities lost function_calling, response_schema, tool_choice | 2026-09-18 |
| truncation | ~deepseek | -67% max out · ~deepseek/deepseek-flash-latest: max_output_tokens cut 393,216 to 131,072 +2 aliases | 2026-09-18 |
| truncation | deepseek | -67% max out · deepseek/deepseek-v4-flash: max_output_tokens cut 393,216 to 131,072 +1 alias | 2026-09-24 |
| cost | z-ai | 3.00x price · z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-09-28 |
| cost | deepseek | 2.99x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.05/Mtok to $0.14/Mtok | 2026-09-28 |
| cost | deepseek | 2.99x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.09/Mtok to $0.28/Mtok | 2026-09-28 |
| cost | deepseek | 2.99x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.0094/Mtok to $0.03/Mtok | 2026-09-28 |
| cost | openrouter | 2.98x price · openrouter/deepseek/deepseek-v4-flash: input_price_per_mtok $0.05/Mtok to $0.14/Mtok | 2026-09-28 |
| cost | openrouter | 2.98x price · openrouter/deepseek/deepseek-v4-flash: output_price_per_mtok $0.09/Mtok to $0.28/Mtok | 2026-09-28 |
| cost | openrouter | 2.98x price · openrouter/deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.0094/Mtok to $0.03/Mtok | 2026-09-28 |
| truncation | arcee-ai | -66% max out · arcee-ai/trinity-large-thinking: max_output_tokens cut 235,929 to 80,000 inferred | 2026-08-30 |
| truncation | ~deepseek | -66% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 384,000 to 131,072 +1 alias | 2026-09-18 |
| truncation | deepseek | -66% max out · deepseek/deepseek-v4-flash-0731: max_output_tokens cut 384,000 to 131,072 +1 alias | 2026-09-27 |
| truncation | qwen | -65% max out · qwen/qwen3.5-122b-a10b: max_output_tokens cut 235,929 to 81,920 +2 aliases | 2026-08-28 |
| cost | openrouter | 2.86x price · openrouter/deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.42/Mtok to $1.20/Mtok | 2026-09-24 |
| cost | deepseek | 2.86x price · deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.42/Mtok to $1.20/Mtok | 2026-09-26 |
| cost | ~deepseek | 2.85x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0080/Mtok to $0.02/Mtok | 2026-09-23 |
| cost | ~deepseek | 2.82x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.09/Mtok to $0.25/Mtok | 2026-09-25 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: output_price_per_mtok $0.10/Mtok to $0.28/Mtok +1 alias | 2026-09-29 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: cache_read_price_per_mtok $0.0100/Mtok to $0.03/Mtok +1 alias | 2026-09-29 |
| cost | deepseek | 2.80x price · deepseek/deepseek-v4-flash-0731: input_price_per_mtok $0.05/Mtok to $0.14/Mtok +1 alias | 2026-09-29 |
| cost | ~deepseek | 2.78x price · ~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.0036/Mtok to $0.01/Mtok | 2026-09-23 |
| cost | deepseek | 2.75x price · deepseek/deepseek-v4-pro: output_price_per_mtok $0.70/Mtok to $1.91/Mtok | 2026-09-28 |
| cost | deepseek | 2.75x price · deepseek/deepseek-v4-pro: input_price_per_mtok $0.35/Mtok to $0.96/Mtok | 2026-09-28 |
| cost | deepseek | 2.74x price · deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.03/Mtok to $0.08/Mtok | 2026-09-28 |
| cost | ~deepseek | 2.74x price · ~deepseek/deepseek-pro-latest: output_price_per_mtok $1.05/Mtok to $2.88/Mtok | 2026-09-25 |
| truncation | nebius | -63% max out · nebius/deepseek-ai/DeepSeek-V4.1-Flash: max_output_tokens cut 1,048,576 to 384,000 | 2026-09-29 |
| cost | openrouter | 2.73x price · openrouter/deepseek/deepseek-v4-pro: output_price_per_mtok $0.70/Mtok to $1.90/Mtok | 2026-09-28 |
| cost | openrouter | 2.73x price · openrouter/deepseek/deepseek-v4-pro: input_price_per_mtok $0.35/Mtok to $0.95/Mtok | 2026-09-28 |
| cost | openrouter | 2.72x price · openrouter/deepseek/deepseek-v4-pro: cache_read_price_per_mtok $0.03/Mtok to $0.08/Mtok | 2026-09-28 |
| cost | ~z-ai | 2.67x price · ~z-ai/glm-latest: output_price_per_mtok $1.50/Mtok to $4.00/Mtok | 2026-09-30 |
| cost | ~moonshotai | 2.62x price · ~moonshotai/kimi-latest: cache_read_price_per_mtok $0.11/Mtok to $0.30/Mtok | 2026-09-27 |
| truncation | novita | -62% max out · novita/moonshotai/kimi-k2-0905: max_output_tokens cut 262,144 to 100,352 +1 alias | 2026-08-28 |
| cost | ~deepseek | 2.59x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok | 2026-09-23 |
| truncation | deepseek | -61% ctx · deepseek/deepseek-r1: context_tokens cut 163,840 to 64,000 | 2026-08-13 |
| cost | ~deepseek | 2.53x price · ~deepseek/deepseek-pro-latest: input_price_per_mtok $0.10/Mtok to $0.25/Mtok | 2026-09-25 |
| cost | ~z-ai | 2.51x price · ~z-ai/glm-latest: output_price_per_mtok $1.12/Mtok to $2.81/Mtok | 2026-09-28 |
| truncation | deepseek | -60% max out · deepseek/deepseek-v3.2: max_output_tokens cut 163,840 to 65,536 +1 alias | 2026-08-09 |
| truncation | novita | -60% max out · novita/deepseek/deepseek-v3-0324: max_output_tokens cut 163,840 to 65,536 inferred | 2026-08-28 |
| truncation | openrouter | -60% max out · openrouter/deepseek/deepseek-v3.2-exp: max_output_tokens cut 163,840 to 65,536 +1 alias | 2026-09-18 |
| cost | openrouter | 2.50x price · openrouter/~deepseek/deepseek-flash-latest: output_price_per_mtok $0.48/Mtok to $1.20/Mtok | 2026-09-23 |
| cost | openrouter | 2.50x price · openrouter/~deepseek/deepseek-flash-latest: input_price_per_mtok $0.12/Mtok to $0.30/Mtok | 2026-09-23 |
| cost | ~deepseek | 2.50x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.50/Mtok to $1.25/Mtok | 2026-09-29 |
| cost | ~deepseek | 2.48x price · ~deepseek/deepseek-flash-latest: input_price_per_mtok $0.04/Mtok to $0.10/Mtok | 2026-09-25 |
| truncation | openrouter | -59% max out · openrouter/~deepseek/deepseek-flash-latest: max_output_tokens cut 943,718 to 384,000 | 2026-09-21 |
| truncation | deepseek | -59% max out · deepseek/deepseek-v4-pro-0813: max_output_tokens cut 943,718 to 384,000 +3 aliases | 2026-09-28 |
| truncation | ~deepseek | -59% max out · ~deepseek/deepseek-flash-latest: max_output_tokens cut 943,718 to 384,000 +3 aliases | 2026-09-28 |
| truncation | deepseek | -59% max out · deepseek/deepseek-v4-pro-0813: max_output_tokens cut 943,717 to 384,000 | 2026-08-28 |
| truncation | novita | -59% max out · novita/qwen/qwen3-4b-fp8: max_output_tokens cut 20,000 to 8,192 inferred | 2026-08-28 |
| cost | ~deepseek | 2.42x price · ~deepseek/deepseek-pro-latest: output_price_per_mtok $1.20/Mtok to $2.90/Mtok | 2026-09-23 |
| truncation | openrouter | -58% max out · openrouter/deepseek/deepseek-v4.1-flash: max_output_tokens cut 943,718 to 393,216 | 2026-09-24 |
| truncation | ~deepseek | -58% max out · ~deepseek/deepseek-flash-latest: max_output_tokens cut 943,718 to 393,216 +10 aliases | 2026-09-29 |
| truncation | deepseek | -58% max out · deepseek/deepseek-v4-pro-0813: max_output_tokens cut 943,718 to 393,216 +1 alias | 2026-09-30 |
| truncation | moonshotai | -57% max out · moonshotai/kimi-k2-thinking: max_output_tokens cut 235,929 to 100,352 +2 aliases | 2026-09-13 |
| cost | z-ai | 2.35x price · z-ai/glm-5.2: input_price_per_mtok $0.21/Mtok to $0.49/Mtok | 2026-09-30 |
| truncation | deepseek | -56% max out · deepseek/deepseek-v3.1-terminus: max_output_tokens cut 147,456 to 65,536 +3 aliases | 2026-09-29 |
| truncation | sao10k | -55% max out · sao10k/l3-lunaris-8b: max_output_tokens cut 16,384 to 7,372 inferred | 2026-08-25 |
| truncation | openrouter | -55% max out · openrouter/qwen/qwen3-235b-a22b-thinking-2507: max_output_tokens cut 262,144 to 117,964 | 2026-09-18 |
| cost | ~z-ai | 2.22x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.07/Mtok to $0.15/Mtok | 2026-09-28 |
| cost | z-ai | 2.16x price · z-ai/glm-5.2: output_price_per_mtok $2.04/Mtok to $4.40/Mtok | 2026-09-29 |
| cost | z-ai | 2.15x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.12/Mtok to $0.26/Mtok | 2026-09-29 |
| cost | openrouter | 2.14x price · openrouter/deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.14/Mtok to $0.30/Mtok | 2026-09-24 |
| cost | deepseek | 2.14x price · deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.14/Mtok to $0.30/Mtok | 2026-09-26 |
| truncation | openai | -53% max out · gpt-5-pro-2025-10-06: max_output_tokens cut 272,000 to 128,000 +1 alias | 2026-06-26 |
| cost | qwen | 2.08x price · qwen/qwen3-30b-a3b-instruct-2507: input_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-09-24 |
| cost | openrouter | 2.08x price · openrouter/qwen/qwen3-30b-a3b-instruct-2507: input_price_per_mtok $0.05/Mtok to $0.10/Mtok | 2026-09-24 |
| cost | openrouter | 2.07x price · openrouter/deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.29/Mtok to $0.60/Mtok | 2026-09-28 |
| cost | deepseek | 2.07x price · deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.29/Mtok to $0.60/Mtok +1 alias | 2026-09-28 |
| cost | ~deepseek | 2.07x price · ~deepseek/deepseek-flash-latest: output_price_per_mtok $0.29/Mtok to $0.60/Mtok +2 aliases | 2026-09-29 |
| truncation | z-ai | -51% max out · z-ai/glm-5.2: max_output_tokens cut 262,144 to 128,000 | 2026-08-06 |
| cost | deepseek | 2.04x price · deepseek/deepseek-v4-flash-vision-exp: output_price_per_mtok $0.65/Mtok to $1.32/Mtok | 2026-09-28 |
| cost | deepseek | 2.04x price · deepseek/deepseek-v4-flash-vision-exp: cache_read_price_per_mtok $0.0069/Mtok to $0.01/Mtok | 2026-09-28 |
| cost | deepseek | 2.04x price · deepseek/deepseek-v4-flash-vision-exp: input_price_per_mtok $0.22/Mtok to $0.44/Mtok | 2026-09-28 |
| cost | deepseek | 2.03x price · deepseek/deepseek-v3.2-exp: output_price_per_mtok $0.20/Mtok to $0.41/Mtok | 2026-09-29 |
| cost | deepseek | 2.01x price · deepseek/deepseek-v3.2-exp: input_price_per_mtok $0.13/Mtok to $0.27/Mtok | 2026-09-29 |
| cost | ~deepseek | 2.00x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.07/Mtok to $0.15/Mtok | 2026-09-29 |
| cost | openrouter | 2.00x price · openrouter/~moonshotai/kimi-latest: input_price_per_mtok $1.50/Mtok to $3.00/Mtok | 2026-09-23 |
| capability | vertex_ai-language-models | lost code_execution, file_search · vertex_ai/gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search | 2026-06-27 |
| capability | openrouter | lost code_execution, file_search · openrouter/google/gemini-3.1-flash-lite-preview: capabilities lost code_execution, file_search +1 alias | 2026-06-27 |
| truncation | xai | -50% max out · xai/grok-4.20-0309-reasoning: max_output_tokens cut 2,000,000 to 1,000,000 +3 aliases inferred | 2026-08-12 |
| truncation | xai | -50% ctx · xai/grok-4.20-0309-reasoning: context_tokens cut 2,000,000 to 1,000,000 +3 aliases inferred | 2026-08-12 |
| truncation | nvidia | -50% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 262,144 to 131,072 | 2026-08-16 |
| truncation | qwen | -50% max out · qwen/qwen3.6-27b: max_output_tokens cut 131,072 to 65,536 | 2026-08-19 |
| truncation | mistral | -50% max out · mistral/codestral-2508: max_output_tokens cut 256,000 to 128,000 | 2026-08-21 |
| truncation | mistral | -50% ctx · mistral/codestral-2508: context_tokens cut 256,000 to 128,000 | 2026-08-21 |
| capability | qwen | lost logit_bias, min_p · qwen/qwen3-235b-a22b-thinking-2507: capabilities lost logit_bias, min_p | 2026-08-24 |
| truncation | qwen | -50% ctx · qwen/qwen3-235b-a22b-thinking-2507: context_tokens cut 262,144 to 131,072 | 2026-08-24 |
| truncation | -50% ctx · google/gemma-3-27b-it: context_tokens cut 262,144 to 131,072 | 2026-08-28 | |
| capability | xai | lost function_calling, tool_choice · xai/grok-4.20-multi-agent-0309: capabilities lost function_calling, tool_choice inferred | 2026-08-30 |
| capability | z-ai | lost logprobs, top_logprobs · z-ai/glm-4.7: capabilities lost logprobs, top_logprobs | 2026-08-31 |
| capability | kwaipilot | lost logprobs, top_logprobs · kwaipilot/kat-coder-pro-v2.5: capabilities lost logprobs, top_logprobs +1 alias inferred | 2026-08-31 |
| capability | qwen | lost logprobs, top_logprobs · qwen/qwen3-32b: capabilities lost logprobs, top_logprobs | 2026-09-01 |
| truncation | gryphe | -50% max out · gryphe/mythomax-l2-13b: max_output_tokens cut 7,372 to 3,686 | 2026-09-01 |
| capability | xiaomi | lost logprobs, top_logprobs · xiaomi/mimo-v2.5: capabilities lost logprobs, top_logprobs | 2026-09-02 |
| truncation | z-ai | -50% max out · z-ai/glm-5.2: max_output_tokens cut 262,144 to 131,072 +7 aliases | 2026-09-02 |
| capability | meta-llama | lost logprobs, top_logprobs · meta-llama/llama-3.1-70b-instruct: capabilities lost logprobs, top_logprobs | 2026-09-04 |
| truncation | watsonx | -50% max out · watsonx/bigscience/mt0-xxl-13b: max_output_tokens cut 8,192 to 4,096 inferred | 2026-09-06 |
| truncation | watsonx | -50% ctx · watsonx/bigscience/mt0-xxl-13b: context_tokens cut 8,192 to 4,096 inferred | 2026-09-06 |
| truncation | qwen | -50% max out · qwen/qwen3.6-27b: max_output_tokens cut 262,144 to 131,072 +5 aliases | 2026-09-09 |
| capability | z-ai | lost logit_bias, min_p · z-ai/glm-5: capabilities lost logit_bias, min_p | 2026-09-10 |
| truncation | openrouter | -50% max out · openrouter/qwen/qwen3.8-2.4t-a95b: max_output_tokens cut 262,144 to 131,072 | 2026-09-18 |
| truncation | openrouter | -50% ctx · openrouter/qwen/qwen3-235b-a22b-thinking-2507: context_tokens cut 262,144 to 131,072 | 2026-09-18 |
| truncation | cohere | -50% ctx · embed-english-light-v3.0: context_tokens cut 1,024 to 512 +3 aliases inferred | 2026-09-18 |
| truncation | meta-llama | -50% max out · meta-llama/llama-3.1-70b-instruct: max_output_tokens cut 16,384 to 8,192 +3 aliases | 2026-09-21 |
| capability | ~x-ai | lost stop, top_k · ~x-ai/grok-latest: capabilities lost stop, top_k | 2026-09-21 |
| capability | z-ai | lost logprobs, top_logprobs · z-ai/glm-5.3-flash:batch: capabilities lost logprobs, top_logprobs; gained min_p, seed +1 alias | 2026-09-22 |
| capability | thinkingmachines | lost logprobs, top_logprobs · thinkingmachines/inkling-small: capabilities lost logprobs, top_logprobs +1 alias | 2026-09-22 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.02/Mtok to $0.04/Mtok +7 aliases | 2026-09-23 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4-pro-0813: output_price_per_mtok $1.98/Mtok to $3.96/Mtok +7 aliases | 2026-09-23 |
| cost | openrouter | 2.00x price · openrouter/z-ai/glm-5.3-flash: input_price_per_mtok $0.07/Mtok to $0.15/Mtok +3 aliases | 2026-09-23 |
| cost | openrouter | 2.00x price · openrouter/~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.0080/Mtok to $0.02/Mtok | 2026-09-23 |
| cost | openrouter | 2.00x price · openrouter/z-ai/glm-5.3-flash: output_price_per_mtok $0.25/Mtok to $0.50/Mtok +3 aliases | 2026-09-23 |
| truncation | qwen | -50% max out · qwen/qwen3-235b-a22b-2507: max_output_tokens cut 32,768 to 16,384 +3 aliases | 2026-09-25 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0030/Mtok to $0.0060/Mtok +8 aliases | 2026-09-25 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.15/Mtok to $0.30/Mtok +8 aliases | 2026-09-25 |
| truncation | thinkingmachines | -50% ctx · thinkingmachines/inkling-small: context_tokens cut 1,048,576 to 524,288 +1 alias | 2026-09-26 |
| truncation | openrouter | -50% max out · openrouter/meta-llama/llama-3.1-70b-instruct: max_output_tokens cut 16,384 to 8,192 +1 alias | 2026-09-27 |
| truncation | openrouter | -50% max out · openrouter/qwen/qwen3-vl-30b-a3b-instruct: max_output_tokens cut 32,768 to 16,384 | 2026-09-27 |
| cost | openrouter | 2.00x price · openrouter/deepseek/deepseek-v4-flash-vision-exp: cache_read_price_per_mtok $0.0070/Mtok to $0.01/Mtok | 2026-09-28 |
| cost | openrouter | 2.00x price · openrouter/deepseek/deepseek-v4-flash-vision-exp: output_price_per_mtok $0.66/Mtok to $1.32/Mtok | 2026-09-28 |
| cost | openrouter | 2.00x price · openrouter/deepseek/deepseek-v4-flash-vision-exp: input_price_per_mtok $0.22/Mtok to $0.44/Mtok | 2026-09-28 |
| cost | openrouter | 2.00x price · openrouter/z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.03/Mtok +1 alias | 2026-09-28 |
| capability | meta | lost logprobs, top_logprobs · meta/muse-glimmer-30b: capabilities lost logprobs, top_logprobs | 2026-09-28 |
| cost | deepseek | 2.00x price · deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.60/Mtok to $1.20/Mtok +9 aliases | 2026-09-28 |
| cost | openai | 2.00x price · openai/gpt-5.6-sol-pro: cache_read_price_per_mtok $0.20/Mtok to $0.40/Mtok | 2026-09-28 |
| cost | openai | 2.00x price · openai/gpt-5.6-sol-pro: input_price_per_mtok $2.00/Mtok to $4.00/Mtok | 2026-09-28 |
| cost | openai | 2.00x price · openai/gpt-5.6-sol-pro: output_price_per_mtok $10.00/Mtok to $20.00/Mtok | 2026-09-28 |
| capability | xiaomi | lost logit_bias, min_p · xiaomi/mimo-v2.5-pro: capabilities lost logit_bias, min_p +1 alias | 2026-09-29 |
| truncation | openrouter | -50% ctx · openrouter/thinkingmachines/inkling-small: context_tokens cut 1,048,576 to 524,288 +1 alias | 2026-09-29 |
| cost | openrouter | 2.00x price · openrouter/openai/gpt-5.6-sol-pro: input_price_per_mtok $2.00/Mtok to $4.00/Mtok | 2026-09-29 |
| cost | openrouter | 2.00x price · openrouter/deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.60/Mtok to $1.20/Mtok +1 alias | 2026-09-29 |
| cost | openrouter | 2.00x price · openrouter/openai/gpt-5.6-sol-pro: cache_read_price_per_mtok $0.20/Mtok to $0.40/Mtok | 2026-09-29 |
| cost | openrouter | 2.00x price · openrouter/openai/gpt-5.6-sol-pro: output_price_per_mtok $10.00/Mtok to $20.00/Mtok | 2026-09-29 |
| truncation | qwen | -50% max out · qwen/qwen3-14b: max_output_tokens cut 16,384 to 8,192 +5 aliases | 2026-09-30 |
| truncation | nvidia | -49% ctx · nvidia/nemotron-3-ultra-550b-a55b: context_tokens cut 512,288 to 262,144 | 2026-08-28 |
| truncation | liquid | -49% ctx · liquid/lfm-2.5-2.6b:free: context_tokens cut 128,000 to 65,536 inferred | 2026-08-21 |
| truncation | mistralai | -49% ctx · mistralai/mistral-small-3.2-24b-instruct: context_tokens cut 256,000 to 131,072 | 2026-08-22 |
| truncation | watsonx | -49% max out · watsonx/mistralai/mistral-small-3-1-24b-instruct-2503: max_output_tokens cut 32,000 to 16,384 inferred | 2026-09-06 |
| cost | openrouter | 1.90x price · openrouter/qwen/qwen3-14b: input_price_per_mtok $0.12/Mtok to $0.23/Mtok +1 alias | 2026-09-28 |
| truncation | ~z-ai | -46% max out · ~z-ai/glm-latest: max_output_tokens cut 235,929 to 128,000 inferred | 2026-09-10 |
| cost | deepseek | 1.83x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-30 |
| cost | deepseek | 1.83x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.08/Mtok to $0.14/Mtok | 2026-09-30 |
| cost | deepseek | 1.83x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.15/Mtok to $0.28/Mtok | 2026-09-30 |
| cost | fireworks_ai | 1.82x price · fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: output_price_per_mtok $0.66/Mtok to $1.20/Mtok +1 alias | 2026-09-26 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-flash: output_price_per_mtok $0.10/Mtok to $0.18/Mtok | 2026-09-23 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-flash: input_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-09-23 |
| cost | openrouter | 1.81x price · openrouter/deepseek/deepseek-v4-flash: input_price_per_mtok $0.05/Mtok to $0.09/Mtok | 2026-09-24 |
| cost | openrouter | 1.81x price · openrouter/deepseek/deepseek-v4-flash: output_price_per_mtok $0.10/Mtok to $0.18/Mtok | 2026-09-24 |
| cost | deepseek | 1.81x price · deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.0098/Mtok to $0.02/Mtok | 2026-09-23 |
| cost | openrouter | 1.81x price · openrouter/deepseek/deepseek-v4-flash: cache_read_price_per_mtok $0.0098/Mtok to $0.02/Mtok | 2026-09-24 |
| truncation | openrouter | -44% max out · openrouter/~z-ai/glm-latest: max_output_tokens cut 235,929 to 131,072 | 2026-09-19 |
| truncation | nvidia | -44% max out · nvidia/nemotron-3.5-lightning: max_output_tokens cut 235,929 to 131,072 +2 aliases | 2026-09-26 |
| truncation | ~z-ai | -44% max out · ~z-ai/glm-latest: max_output_tokens cut 235,929 to 131,072 +2 aliases | 2026-09-27 |
| truncation | qwen | -44% max out · qwen/qwen3.8-27b: max_output_tokens cut 235,929 to 131,072 | 2026-09-30 |
| truncation | openai | -44% max out · openai/gpt-oss-120b: max_output_tokens cut 117,964 to 65,536 | 2026-09-18 |
| truncation | openrouter | -44% max out · openrouter/openai/gpt-oss-120b: max_output_tokens cut 117,964 to 65,536 | 2026-09-19 |
| cost | ~deepseek | 1.79x price · ~deepseek/deepseek-pro-latest: output_price_per_mtok $1.96/Mtok to $3.50/Mtok +1 alias | 2026-09-29 |
| cost | ~deepseek | 1.77x price · ~deepseek/deepseek-pro-latest: input_price_per_mtok $0.13/Mtok to $0.23/Mtok | 2026-09-28 |
| cost | z-ai | 1.71x price · z-ai/glm-5.3: output_price_per_mtok $2.57/Mtok to $4.40/Mtok | 2026-09-27 |
| cost | openrouter | 1.71x price · openrouter/z-ai/glm-5.3: output_price_per_mtok $2.57/Mtok to $4.40/Mtok | 2026-09-29 |
| cost | openrouter | 1.69x price · openrouter/z-ai/glm-5.3: cache_read_price_per_mtok $0.04/Mtok to $0.07/Mtok | 2026-09-28 |
| cost | openrouter | 1.67x price · openrouter/~deepseek/deepseek-flash-latest: cache_read_price_per_mtok $0.0036/Mtok to $0.0060/Mtok | 2026-09-23 |
| cost | ~deepseek | 1.67x price · ~deepseek/deepseek-flash-latest: output_price_per_mtok $0.60/Mtok to $1.00/Mtok | 2026-09-24 |
| cost | z-ai | 1.67x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.16/Mtok to $0.26/Mtok | 2026-09-24 |
| cost | z-ai | 1.67x price · z-ai/glm-5.3: output_price_per_mtok $2.64/Mtok to $4.40/Mtok | 2026-09-24 |
| cost | openrouter | 1.67x price · openrouter/z-ai/glm-5.3: cache_read_price_per_mtok $0.16/Mtok to $0.26/Mtok | 2026-09-24 |
| cost | openrouter | 1.67x price · openrouter/z-ai/glm-5.3: output_price_per_mtok $2.64/Mtok to $4.40/Mtok | 2026-09-24 |
| cost | ~deepseek | 1.67x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.30/Mtok to $0.50/Mtok | 2026-09-29 |
| cost | openrouter | 1.67x price · openrouter/z-ai/glm-5.3: input_price_per_mtok $0.84/Mtok to $1.40/Mtok | 2026-09-24 |
| cost | z-ai | 1.67x price · z-ai/glm-5.3: input_price_per_mtok $0.84/Mtok to $1.40/Mtok | 2026-09-24 |
| cost | ~moonshotai | 1.64x price · ~moonshotai/kimi-latest: cache_read_price_per_mtok $0.18/Mtok to $0.29/Mtok | 2026-09-28 |
| cost | ~moonshotai | 1.63x price · ~moonshotai/kimi-latest: output_price_per_mtok $5.53/Mtok to $9.00/Mtok | 2026-09-27 |
| cost | tencent | 1.60x price · tencent/hy3: cache_read_price_per_mtok $0.02/Mtok to $0.03/Mtok +34 aliases | 2026-09-30 |
| cost | tencent | 1.60x price · tencent/hy3: output_price_per_mtok $0.33/Mtok to $0.53/Mtok +34 aliases | 2026-09-30 |
| cost | tencent | 1.60x price · tencent/hy3: input_price_per_mtok $0.08/Mtok to $0.13/Mtok +34 aliases | 2026-09-30 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.25/Mtok to $0.40/Mtok | 2026-09-28 |
| cost | deepseek | 1.58x price · deepseek/deepseek-v4-pro-0813: input_price_per_mtok $0.25/Mtok to $0.40/Mtok | 2026-09-28 |
| truncation | openrouter | -36% max out · openrouter/minimax/minimax-m2: max_output_tokens cut 204,800 to 131,072 | 2026-09-18 |
| cost | z-ai | 1.56x price · z-ai/glm-5.2: cache_read_price_per_mtok $0.17/Mtok to $0.26/Mtok | 2026-09-30 |
| cost | openrouter | 1.55x price · openrouter/qwen/qwen3-30b-a3b-instruct-2507: output_price_per_mtok $0.19/Mtok to $0.30/Mtok +1 alias | 2026-09-24 |
| cost | qwen | 1.55x price · qwen/qwen3-30b-a3b-instruct-2507: output_price_per_mtok $0.19/Mtok to $0.30/Mtok +3 aliases | 2026-09-24 |
| truncation | fireworks_ai | -35% max out · fireworks_ai/accounts/fireworks/models/glm-5p1: max_output_tokens cut 202,800 to 131,072 +1 alias inferred | 2026-06-18 |
| cost | ~moonshotai | 1.54x price · ~moonshotai/kimi-latest: output_price_per_mtok $6.99/Mtok to $10.75/Mtok | 2026-09-24 |
| cost | ~moonshotai | 1.54x price · ~moonshotai/kimi-latest: cache_read_price_per_mtok $0.20/Mtok to $0.30/Mtok | 2026-09-24 |
| cost | ~moonshotai | 1.54x price · ~moonshotai/kimi-latest: input_price_per_mtok $0.91/Mtok to $1.40/Mtok | 2026-09-24 |
| cost | ~z-ai | 1.53x price · ~z-ai/glm-latest: input_price_per_mtok $0.18/Mtok to $0.27/Mtok | 2026-09-27 |
| cost | z-ai | 1.53x price · z-ai/glm-5.3: input_price_per_mtok $0.18/Mtok to $0.27/Mtok | 2026-09-27 |
| cost | z-ai | 1.53x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.03/Mtok to $0.04/Mtok | 2026-09-27 |
| cost | ~z-ai | 1.53x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.03/Mtok to $0.04/Mtok | 2026-09-27 |
| truncation | undi95 | -33% max out · undi95/remm-slerp-l2-13b: max_output_tokens cut 6,144 to 4,096 | 2026-08-24 |
| cost | z-ai | 1.50x price · z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.01/Mtok | 2026-09-26 |
| cost | ~z-ai | 1.50x price · ~z-ai/glm-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.01/Mtok | 2026-09-26 |
| cost | openrouter | 1.50x price · openrouter/z-ai/glm-5.3-flash: cache_read_price_per_mtok $0.01/Mtok to $0.01/Mtok | 2026-09-27 |
| cost | ~deepseek | 1.50x price · ~deepseek/deepseek-flash-latest: input_price_per_mtok $0.02/Mtok to $0.03/Mtok | 2026-09-28 |
| cost | deepseek | 1.50x price · deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.10/Mtok to $0.15/Mtok | 2026-09-23 |
| cost | z-ai | 1.50x price · z-ai/glm-4.7: input_price_per_mtok $0.40/Mtok to $0.60/Mtok | 2026-09-25 |
| cost | openrouter | 1.50x price · openrouter/z-ai/glm-4.7: input_price_per_mtok $0.40/Mtok to $0.60/Mtok | 2026-09-26 |
| cost | ~z-ai | 1.50x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.10/Mtok to $0.15/Mtok | 2026-09-29 |
| cost | openrouter | 1.49x price · openrouter/z-ai/glm-5.3: input_price_per_mtok $0.24/Mtok to $0.36/Mtok | 2026-09-28 |
| truncation | azure_ai | -32% ctx · azure_ai/gpt-5.4-mini-2026-03-17: context_tokens cut 400,000 to 272,000 +3 aliases | 2026-07-30 |
| cost | qwen | 1.47x price · qwen/qwen3.8-27b: output_price_per_mtok $3.00/Mtok to $4.40/Mtok | 2026-09-29 |
| truncation | ~deepseek | -32% max out · ~deepseek/deepseek-v4-flash-latest: max_output_tokens cut 384,000 to 262,144 +3 aliases inferred | 2026-08-19 |
| cost | ~deepseek | 1.46x price · ~deepseek/deepseek-pro-latest: cache_read_price_per_mtok $0.17/Mtok to $0.25/Mtok | 2026-09-28 |
| cost | deepseek | 1.46x price · deepseek/deepseek-v4-pro-0813: cache_read_price_per_mtok $0.17/Mtok to $0.25/Mtok | 2026-09-28 |
| cost | openrouter | 1.45x price · openrouter/z-ai/glm-5.1: cache_read_price_per_mtok $0.18/Mtok to $0.26/Mtok | 2026-09-28 |
| cost | openrouter | 1.45x price · openrouter/z-ai/glm-5.1: output_price_per_mtok $3.03/Mtok to $4.40/Mtok | 2026-09-28 |
| cost | z-ai | 1.45x price · z-ai/glm-5.1: output_price_per_mtok $3.03/Mtok to $4.40/Mtok +2 aliases | 2026-09-30 |
| cost | z-ai | 1.45x price · z-ai/glm-5.1: cache_read_price_per_mtok $0.18/Mtok to $0.26/Mtok +2 aliases | 2026-09-30 |
| cost | openrouter | 1.45x price · openrouter/z-ai/glm-5.1: input_price_per_mtok $0.96/Mtok to $1.40/Mtok | 2026-09-28 |
| cost | z-ai | 1.45x price · z-ai/glm-5.1: input_price_per_mtok $0.96/Mtok to $1.40/Mtok +2 aliases | 2026-09-30 |
| cost | deepseek | 1.43x price · deepseek/deepseek-v4.1-flash: output_price_per_mtok $0.42/Mtok to $0.60/Mtok | 2026-09-24 |
| cost | openrouter | 1.43x price · openrouter/deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0042/Mtok to $0.0060/Mtok | 2026-09-24 |
| cost | deepseek | 1.43x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0042/Mtok to $0.0060/Mtok | 2026-09-26 |
| cost | openrouter | 1.43x price · openrouter/minimax/minimax-m2.7: input_price_per_mtok $0.21/Mtok to $0.30/Mtok | 2026-09-28 |
| cost | minimax | 1.43x price · minimax/minimax-m2.7: input_price_per_mtok $0.21/Mtok to $0.30/Mtok | 2026-09-28 |
| cost | openrouter | 1.43x price · openrouter/minimax/minimax-m2.7: output_price_per_mtok $0.84/Mtok to $1.20/Mtok | 2026-09-28 |
| cost | minimax | 1.43x price · minimax/minimax-m2.7: output_price_per_mtok $0.84/Mtok to $1.20/Mtok | 2026-09-28 |
| cost | openrouter | 1.43x price · openrouter/minimax/minimax-m2.7: cache_read_price_per_mtok $0.04/Mtok to $0.06/Mtok | 2026-09-28 |
| cost | minimax | 1.43x price · minimax/minimax-m2.7: cache_read_price_per_mtok $0.04/Mtok to $0.06/Mtok | 2026-09-28 |
| truncation | z-ai | -30% max out · z-ai/glm-5.1: max_output_tokens cut 182,476 to 128,000 | 2026-08-30 |
| cost | moonshotai | 1.42x price · moonshotai/kimi-k3: output_price_per_mtok $10.53/Mtok to $15.00/Mtok | 2026-09-25 |
| cost | ~z-ai | 1.42x price · ~z-ai/glm-latest: output_price_per_mtok $1.76/Mtok to $2.50/Mtok | 2026-09-24 |
| cost | deepseek | 1.41x price · deepseek/deepseek-v4.1-flash: input_price_per_mtok $0.10/Mtok to $0.14/Mtok | 2026-09-26 |
| cost | deepseek | 1.40x price · deepseek/deepseek-v4.1-flash: cache_read_price_per_mtok $0.0030/Mtok to $0.0042/Mtok | 2026-09-24 |
| cost | openrouter | 1.39x price · openrouter/~moonshotai/kimi-latest: output_price_per_mtok $10.76/Mtok to $15.00/Mtok | 2026-09-23 |
| truncation | z-ai | -28% max out · z-ai/glm-5.2: max_output_tokens cut 182,476 to 131,072 | 2026-09-14 |
| cost | z-ai | 1.38x price · z-ai/glm-4.6: cache_read_price_per_mtok $0.08/Mtok to $0.11/Mtok +2 aliases | 2026-09-25 |
| cost | openrouter | 1.38x price · openrouter/z-ai/glm-4.7: cache_read_price_per_mtok $0.08/Mtok to $0.11/Mtok | 2026-09-26 |
| cost | z-ai | 1.37x price · z-ai/glm-5.3: cache_read_price_per_mtok $0.19/Mtok to $0.26/Mtok | 2026-09-29 |
| cost | fireworks_ai | 1.36x price · fireworks_ai/accounts/fireworks/routers/kimi-k3-us: input_price_per_mtok $3.30/Mtok to $4.50/Mtok +1 alias | 2026-09-24 |
| cost | fireworks_ai | 1.36x price · fireworks_ai/accounts/fireworks/routers/kimi-k3-us: cache_read_price_per_mtok $0.33/Mtok to $0.45/Mtok +1 alias | 2026-09-24 |
| cost | fireworks_ai | 1.36x price · fireworks_ai/accounts/fireworks/routers/kimi-k3-us: output_price_per_mtok $16.50/Mtok to $22.50/Mtok +1 alias | 2026-09-24 |
| cost | fireworks_ai | 1.36x price · fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: input_price_per_mtok $0.22/Mtok to $0.30/Mtok +1 alias | 2026-09-26 |
| cost | ~moonshotai | 1.36x price · ~moonshotai/kimi-latest: input_price_per_mtok $0.88/Mtok to $1.20/Mtok | 2026-09-25 |
| truncation | minimax | -26% max out · minimax/minimax-m2.7: max_output_tokens cut 176,947 to 131,072 +2 aliases | 2026-09-28 |
| cost | openrouter | 1.33x price · openrouter/~deepseek/deepseek-v4-flash-latest: input_price_per_mtok $0.03/Mtok to $0.04/Mtok | 2026-09-23 |
| cost | 1.33x price · google/gemma-4-26b-a4b-it: cache_read_price_per_mtok $0.04/Mtok to $0.05/Mtok | 2026-09-28 | |
| cost | ~moonshotai | 1.33x price · ~moonshotai/kimi-latest: cache_read_price_per_mtok $0.30/Mtok to $0.40/Mtok | 2026-09-29 |
| cost | ~deepseek | 1.33x price · ~deepseek/deepseek-v4-flash-latest: cache_read_price_per_mtok $0.01/Mtok to $0.02/Mtok | 2026-09-26 |
| cost | 1.33x price · google/gemma-4-26b-a4b-it: input_price_per_mtok $0.07/Mtok to $0.09/Mtok | 2026-09-28 | |
| cost | 1.33x price · google/gemma-4-26b-a4b-it: output_price_per_mtok $0.23/Mtok to $0.30/Mtok | 2026-09-28 | |
| cost | ~z-ai | 1.32x price · ~z-ai/glm-latest: cache_read_price_per_mtok $0.04/Mtok to $0.06/Mtok | 2026-09-27 |
| cost | ~deepseek | 1.31x price · ~deepseek/deepseek-v4-flash-latest: output_price_per_mtok $0.32/Mtok to $0.42/Mtok | 2026-09-25 |
| cost | ~moonshotai | 1.31x price · ~moonshotai/kimi-latest: output_price_per_mtok $8.54/Mtok to $11.20/Mtok | 2026-09-28 |
| truncation | novita | -23% max out · novita/moonshotai/kimi-k2-instruct: max_output_tokens cut 131,072 to 100,352 inferred | 2026-08-28 |
Announced deprecations
Changes that came with a published date. Separated out because they are a different problem: you were told.
| Impact | Provider | What changed | Detected |
|---|---|---|---|
| availability | qwen | qwen/qwen-plus-2025-07-28: retirement announced for 2026-10-09 — in 14 days +15 aliases | 2026-09-30 |
| availability | qwen | qwen/qwen-plus-2025-07-28: retires 2026-10-09 (in 14 days) +14 aliases | 2026-09-29 |
| availability | openrouter | openrouter/qwen/qwen-plus-2025-07-28: retires 2026-10-09 (in 13 days) +15 aliases | 2026-09-28 |
| availability | openrouter | openrouter/qwen/qwen-plus-2025-07-28: deprecation announced for 2026-10-09 — in 13 days +15 aliases | 2026-09-28 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: deprecation announced for 2026-09-15 (date already passed) +5 aliases | 2026-09-28 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Pro: deprecation announced for 2026-08-27 (date already passed) +4 aliases | 2026-09-25 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: deprecation announced for 2026-09-14 (date already passed) +5 aliases | 2026-09-25 |
| availability | antigravity-preview-09-2026: retires 2026-09-01 (23 days ago) inferred | 2026-09-24 | |
| availability | azure_ai | azure_ai/Cohere-command-a-plus-05-2026: deprecation announced for 2026-10-16 — in 22 days inferred | 2026-09-24 |
| availability | azure_ai | azure_ai/Cohere-command-a-plus-05-2026: retires 2026-10-16 (in 22 days) inferred | 2026-09-24 |
| availability | azure | azure/eu/gpt-4.1-nano: deprecation announced for 2026-10-14 — in 20 days +7 aliases | 2026-09-24 |
| availability | vertex_ai-moonshot_models | vertex_ai/moonshotai/kimi-k2-thinking-maas: retires 2026-10-21 (in 27 days) inferred | 2026-09-24 |
| availability | vertex_ai-deepseek_models | vertex_ai/deepseek-ai/deepseek-r1-0528-maas: retires 2026-10-21 (in 27 days) +2 aliases inferred | 2026-09-24 |
| availability | vertex_ai-zai_models | vertex_ai/zai-org/glm-4.7-maas: deprecation announced for 2026-10-21 — in 27 days +1 alias inferred | 2026-09-24 |
| availability | fireworks_ai | fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: deprecation announced for 2026-09-25 — in 1 days +17 aliases | 2026-09-24 |
| availability | vertex_ai-qwen_models | vertex_ai/qwen/qwen3-235b-a22b-instruct-2507-maas: deprecation announced for 2026-10-21 — in 27 days +3 aliases inferred | 2026-09-24 |
| availability | vertex_ai-deepseek_models | vertex_ai/deepseek-ai/deepseek-r1-0528-maas: deprecation announced for 2026-10-21 — in 27 days +2 aliases inferred | 2026-09-24 |
| availability | vertex_ai | vertex_ai/deepseek-ai/deepseek-ocr-maas: retires 2026-10-21 (in 27 days) inferred | 2026-09-24 |
| availability | vertex_ai-openai_models | vertex_ai/openai/gpt-oss-20b-maas: deprecation announced for 2026-10-21 — in 27 days inferred | 2026-09-24 |
| availability | fireworks_ai | fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: retires 2026-09-25 (in 1 days) +17 aliases | 2026-09-24 |
| availability | vertex_ai-minimax_models | vertex_ai/minimaxai/minimax-m2-maas: deprecation announced for 2026-10-21 — in 27 days inferred | 2026-09-24 |
| availability | vertex_ai-minimax_models | vertex_ai/minimaxai/minimax-m2-maas: retires 2026-10-21 (in 27 days) inferred | 2026-09-24 |
| availability | vertex_ai-zai_models | vertex_ai/zai-org/glm-4.7-maas: retires 2026-10-21 (in 27 days) +1 alias inferred | 2026-09-24 |
| availability | vertex_ai-moonshot_models | vertex_ai/moonshotai/kimi-k2-thinking-maas: deprecation announced for 2026-10-21 — in 27 days inferred | 2026-09-24 |
| availability | vertex_ai-llama_models | vertex_ai/meta/llama-3.3-70b-instruct-maas: deprecation announced for 2026-10-21 — in 27 days inferred | 2026-09-24 |
| availability | vertex_ai-llama_models | vertex_ai/meta/llama-3.3-70b-instruct-maas: retires 2026-10-21 (in 27 days) inferred | 2026-09-24 |
| availability | vertex_ai | vertex_ai/deepseek-ai/deepseek-ocr-maas: deprecation announced for 2026-10-21 — in 27 days inferred | 2026-09-24 |
| availability | vertex_ai-openai_models | vertex_ai/openai/gpt-oss-20b-maas: retires 2026-10-21 (in 27 days) inferred | 2026-09-24 |
| availability | vertex_ai-qwen_models | vertex_ai/qwen/qwen3-235b-a22b-instruct-2507-maas: retires 2026-10-21 (in 27 days) +3 aliases inferred | 2026-09-24 |
| availability | fireworks_ai | fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: retires 2026-08-27 (20 days ago) +7 aliases | 2026-09-24 |
| availability | fireworks_ai | fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: deprecation announced for 2026-08-27 (date already passed) +7 aliases | 2026-09-24 |
| availability | openai | openai/sora-2-pro-high-res: retires 2026-09-24 (in 0 days) +8 aliases | 2026-09-24 |
| availability | openrouter | openrouter/nex-agi/nex-n2.5-mini:free: retires 2026-09-25 (in 3 days) +1 alias | 2026-09-22 |
| availability | openrouter | openrouter/nex-agi/nex-n2.5-mini:free: deprecation announced for 2026-09-25 — in 3 days +1 alias | 2026-09-22 |
| availability | nex-agi | nex-agi/nex-n2.5-mini:free: retirement announced for 2026-09-25 — in 3 days +1 alias | 2026-09-22 |
| availability | nex-agi | nex-agi/nex-n2.5-mini:free: retires 2026-09-25 (in 3 days) +1 alias | 2026-09-22 |
| availability | azure_ai | azure_ai/MAI-Image-2.5-Flash: retires 2026-10-01 (in 28 days) +2 aliases inferred | 2026-09-21 |
| availability | azure_ai | azure_ai/MAI-Image-2.5-Flash: deprecation announced for 2026-10-01 — in 28 days +2 aliases inferred | 2026-09-21 |
| availability | groq | groq/qwen/qwen3.6-27b: retires 2026-09-14 (7 days ago) | 2026-09-21 |
| availability | groq | groq/qwen/qwen3.6-27b: deprecation announced for 2026-09-14 (date already passed) | 2026-09-21 |
| availability | openrouter | openrouter/baidu/ernie-4.5-vl-424b-a47b: deprecation announced for 2026-10-08 — in 18 days +1 alias | 2026-09-20 |
| availability | openrouter | openrouter/baidu/ernie-4.5-vl-424b-a47b: retires 2026-10-08 (in 18 days) +1 alias | 2026-09-20 |
| availability | openrouter | openrouter/deepseek/deepseek-r1-distill-llama-70b: deprecation announced for 2026-09-28 — in 8 days +3 aliases | 2026-09-20 |
| availability | openrouter | openrouter/deepseek/deepseek-r1-distill-llama-70b: retires 2026-09-28 (in 8 days) +3 aliases | 2026-09-20 |
| availability | minimax | minimax/minimax-m2.1: retires 2026-10-08 (in 18 days) | 2026-09-20 |
| availability | baidu | baidu/ernie-4.5-vl-424b-a47b: retirement announced for 2026-10-08 — in 18 days | 2026-09-20 |
| availability | deepseek | deepseek/deepseek-r1-distill-llama-70b: retirement announced for 2026-09-28 — in 8 days +3 aliases | 2026-09-20 |
| availability | deepseek | deepseek/deepseek-r1-distill-llama-70b: retires 2026-09-28 (in 8 days) +3 aliases | 2026-09-20 |
| availability | baidu | baidu/ernie-4.5-vl-424b-a47b: retires 2026-10-08 (in 18 days) | 2026-09-20 |
| availability | minimax | minimax/minimax-m2.1: retirement announced for 2026-10-08 — in 18 days | 2026-09-20 |
| availability | databricks | databricks/databricks-gemini-2-5-flash: deprecation announced for 2026-10-02 — in 26 days +1 alias inferred | 2026-09-19 |
| availability | databricks | databricks/databricks-gemini-2-5-flash: retires 2026-10-02 (in 26 days) +1 alias inferred | 2026-09-19 |
| availability | openrouter | openrouter/dots-studio/dots-3-note-preview:free: retires 2026-09-30 (in 12 days) | 2026-09-18 |
| availability | azure | azure/eu/gpt-4o-2024-05-13: retires 2026-10-01 (in 13 days) +7 aliases | 2026-09-18 |
| availability | azure | azure/eu/gpt-4.1-nano: retires 2026-10-14 (in 26 days) +4 aliases | 2026-09-18 |
| availability | azure | azure/eu/gpt-4.1-nano: deprecation announced for 2026-10-14 — in 26 days +2 aliases | 2026-09-18 |
| availability | azure | azure/eu/gpt-4o-2024-05-13: deprecation announced for 2026-10-01 — in 13 days +6 aliases | 2026-09-18 |
| availability | wandb | wandb/MiniMaxAI/MiniMax-M2.5: deprecation announced for 2026-08-25 (date already passed) +1 alias | 2026-09-17 |
| availability | wandb | wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: retires 2026-10-05 (in 18 days) +8 aliases | 2026-09-17 |
| availability | wandb | wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: deprecation announced for 2026-10-05 — in 18 days +8 aliases | 2026-09-17 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Flash-0731: retires 2026-09-29 (in 12 days) +1 alias | 2026-09-17 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Flash-0731: deprecation announced for 2026-09-29 — in 12 days +1 alias | 2026-09-17 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: deprecation announced for 2026-09-14 (date already passed) +3 aliases | 2026-09-17 |
| availability | wandb | wandb/MiniMaxAI/MiniMax-M2.5: retires 2026-08-25 (23 days ago) +1 alias | 2026-09-17 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: retires 2026-09-15 (1 days ago) +3 aliases | 2026-09-16 |
| availability | friendliai | friendliai/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B: retires 2026-09-06 (7 days ago) inferred | 2026-09-13 |
| availability | together_ai | together_ai/moonshotai/Kimi-K2.6: retires 2026-08-19 (25 days ago) | 2026-09-13 |
| availability | together_ai | together_ai/google/gemma-4-31B-it: retires 2026-09-14 (in 1 days) +3 aliases | 2026-09-13 |
| availability | openai | gpt-5.4-cyber: retires 2026-10-01 (in 19 days) inferred | 2026-09-12 |
| availability | dots-studio | dots-studio/dots-3-note-preview:free: retirement announced for 2026-09-30 — in 19 days +1 alias inferred | 2026-09-11 |
| availability | scaleway | scaleway/mistralai/pixtral-12b-2409: retires 2026-10-01 (in 21 days) +1 alias | 2026-09-10 |
| availability | scaleway | scaleway/mistralai/pixtral-12b-2409: deprecation announced for 2026-10-01 — in 21 days +1 alias | 2026-09-10 |
| availability | bedrock_converse | us-gov.anthropic.claude-3-haiku-20240307-v1:0: retires 2026-09-10 (in 1 days) inferred | 2026-09-09 |
| availability | azure | azure/gpt-realtime-2: retires 2026-08-31 (6 days ago) inferred | 2026-09-06 |
| availability | z-ai | z-ai/glm-4.7-flash: retirement announced for 2026-09-10 — in 6 days | 2026-09-04 |
| availability | z-ai | z-ai/glm-4.7-flash: retires 2026-09-10 (in 6 days) | 2026-09-04 |
| availability | cerebras | cerebras/zai-glm-4.7: retires 2026-08-17 (15 days ago) inferred | 2026-09-01 |
| availability | bedrock | amazon.nova-sonic-v1:0: retires 2026-09-14 (in 13 days) inferred | 2026-09-01 |
| availability | cerebras | cerebras/zai-glm-4.7: deprecation announced for 2026-08-17 (date already passed) inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-mini: retires 2026-09-08 (in 7 days) +1 alias inferred | 2026-09-01 |
| availability | nex-agi | nex-agi/nex-n2-mini: retirement announced for 2026-09-08 — in 7 days +1 alias inferred | 2026-09-01 |
| availability | together_ai | together_ai/google/gemma-3n-E4B-it: deprecation announced for 2026-08-25 (date already passed) +1 alias | 2026-08-28 |
| availability | together_ai | together_ai/deepseek-ai/DeepSeek-V4-Pro: retires 2026-08-27 (1 days ago) +3 aliases | 2026-08-28 |
| availability | together_ai | together_ai/google/gemma-3n-E4B-it: retires 2026-08-25 (3 days ago) +1 alias | 2026-08-28 |
| availability | moonshotai | moonshotai/kimi-k2.5: retirement announced for 2026-08-31 — in 4 days +1 alias | 2026-08-27 |
| availability | moonshotai | moonshotai/kimi-k2.5: retires 2026-08-31 (in 7 days) | 2026-08-24 |
| availability | azure_ai | azure_ai/deepseek-r1: deprecation announced for 2026-08-13 (date already passed) | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) inferred | 2026-08-21 |
| availability | gemini | gemini/gemini-robotics-er-1.6-preview: retires 2026-08-31 (in 10 days) | 2026-08-21 |
| availability | azure_ai | azure_ai/deepseek-r1: retires 2026-08-13 (8 days ago) | 2026-08-21 |
| availability | azure_ai | azure_ai/MAI-Image-2e: deprecation announced for 2026-08-15 (date already passed) inferred | 2026-08-21 |
| availability | gemini | gemini/gemini-robotics-er-1.6-preview: deprecation announced for 2026-08-31 — in 10 days | 2026-08-21 |
| availability | azure_ai | azure_ai/MAI-Image-2e: retires 2026-08-15 (6 days ago) inferred | 2026-08-21 |
| availability | azure_ai | azure_ai/claude-opus-4-1: retires 2026-08-05 (16 days ago) inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) +1 alias inferred | 2026-08-21 |
| availability | vertex_ai-anthropic_models | vertex_ai/claude-opus-4-1: retires 2026-08-05 (16 days ago) +1 alias inferred | 2026-08-21 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free: retires 2026-08-24 (in 4 days) +2 aliases inferred | 2026-08-20 |
| availability | nvidia | nvidia/nemotron-3-nano-30b-a3b:free: retirement announced for 2026-08-24 — in 4 days +2 aliases inferred | 2026-08-20 |
| availability | inclusionai | inclusionai/ling-2.6-1t: retires 2026-08-24 (in 5 days) +2 aliases inferred | 2026-08-19 |
| availability | inclusionai | inclusionai/ling-2.6-1t: retirement announced for 2026-08-24 — in 5 days +2 aliases inferred | 2026-08-19 |
| availability | deepseek | deepseek/deepseek-v3.1-terminus: retires 2026-08-17 (in 2 days) | 2026-08-15 |
| availability | deepseek | deepseek/deepseek-v3.1-terminus: retirement announced for 2026-08-17 — in 2 days | 2026-08-15 |
| availability | groq | groq/llama-3.1-8b-instant: deprecation announced for 2026-08-16 — in 3 days +1 alias inferred | 2026-08-13 |
| availability | groq | groq/meta-llama/llama-4-scout-17b-16e-instruct: deprecation announced for 2026-07-17 (date already passed) +1 alias | 2026-08-13 |
| availability | groq | groq/meta-llama/llama-4-scout-17b-16e-instruct: retires 2026-07-17 (27 days ago) +1 alias | 2026-08-13 |
| availability | groq | groq/llama-3.1-8b-instant: retires 2026-08-16 (in 3 days) +1 alias inferred | 2026-08-13 |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free: retirement announced for 2026-08-13 — in 1 days inferred | 2026-08-12 |
| availability | inclusionai | inclusionai/ling-3.0-tiny:free: retires 2026-08-13 (in 1 days) inferred | 2026-08-12 |
| availability | mistral | mistral/devstral-2512: retires 2026-07-31 (12 days ago) +5 aliases inferred | 2026-08-12 |
| availability | anthropic | claude-opus-4-1: deprecation announced for 2026-08-05 (date already passed) inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-sonnet-20240229-v1:0: retires 2026-07-30 (13 days ago) +8 aliases | 2026-08-12 |
| availability | gemini | gemini/imagen-4.0-fast-generate-001: retires 2026-08-17 (in 5 days) +2 aliases | 2026-08-12 |
| availability | bedrock | cohere.command-r-plus-v1:0: retires 2026-08-19 (in 7 days) +1 alias inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-haiku-20240307-v1:0: deprecation announced for 2026-09-10 — in 29 days +5 aliases | 2026-08-12 |
| availability | mistral | mistral/mistral-medium-2505: deprecation announced for 2026-08-31 — in 19 days +2 aliases inferred | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-haiku-20240307-v1:0: retires 2026-09-10 (in 29 days) +5 aliases | 2026-08-12 |
| availability | gemini | gemini/imagen-4.0-fast-generate-001: deprecation announced for 2026-08-17 — in 5 days +2 aliases | 2026-08-12 |
| availability | bedrock | anthropic.claude-3-sonnet-20240229-v1:0: deprecation announced for 2026-07-30 (date already passed) +8 aliases | 2026-08-12 |
| availability | mistral | mistral/mistral-medium-2505: retires 2026-08-31 (in 19 days) +2 aliases inferred | 2026-08-12 |
| availability | bedrock | cohere.command-r-plus-v1:0: deprecation announced for 2026-08-19 — in 7 days +1 alias inferred | 2026-08-12 |
| availability | anthropic | claude-opus-4-1-20250805: retires 2026-08-05 (in 4 days) +1 alias inferred | 2026-08-12 |
| availability | mistral | mistral/devstral-2512: deprecation announced for 2026-07-31 (date already passed) +5 aliases inferred | 2026-08-12 |
| availability | gemini | gemini/gemini-embedding-2-preview: retires 2026-08-10 (2 days ago) | 2026-08-12 |
| availability | gemini | gemini/gemini-embedding-2-preview: deprecation announced for 2026-08-10 (date already passed) | 2026-08-12 |
| availability | openai | gpt-5.2-chat-latest: deprecation announced for 2026-08-10 (date already passed) +1 alias | 2026-08-12 |
| availability | openai | gpt-4o-mini-search-preview-2025-03-11: deprecation announced for 2026-07-23 (date already passed) +15 aliases | 2026-08-12 |
| availability | openai | gpt-4o-audio: deprecation announced for 2026-07-20 (date already passed) +8 aliases inferred | 2026-08-05 |
| availability | openai | computer-use-preview-2025-03-11: retires 2026-07-23 (9 days ago) +17 aliases inferred | 2026-08-01 |
| availability | openai | gpt-5.2-chat-latest: retires 2026-08-10 (in 9 days) +3 aliases | 2026-08-01 |