News Deep Dive Opinion Research Data Resources Events About

Flash Used to Mean Fast and Cheap. DeepSeek's 552B V4.1-Flash and GLM-5.3-Flash Broke the Label

DeepSeek's V4.1-Flash jumps from 284B to 552B parameters and turns reasoning on by default. Zhipu's GLM-5.3-Flash can't switch it off. Here is what happened to the 'fast and cheap' tier.

On September 10, 2026, DeepSeek shipped V4.1-Flash with 552 billion total parameters12. The V4-Flash it replaces had 284 billion3. Two weeks earlier, on August 26, Zhipu released GLM-5.3-Flash: 320 billion parameters, 18 billion active, with a reasoning mode that cannot be turned off45.

Both models carry the name Flash. Neither is small in the way the name suggests, and neither answers quickly by default.

What Flash was supposed to mean

When Google introduced Gemini 1.5 Flash on May 14, 2024, it called the model “lighter-weight than 1.5 Pro” and “designed to be fast and efficient to serve at scale”6. The price matched the pitch: $0.075 per million input tokens, $0.30 per million output tokens7. Flash-8B went lower, at $0.0375 input7.

That was the deal. Flash was never the smartest model in the family. It was the one cheap and quick enough to run at volume.

March 2025: Flash started to think

The turning point came in March 2025, when Google said its Gemini 2.5 models “are thinking models” and that it would build “thinking capabilities directly into all of our models”8.

After that, a Flash model stopped answering on the spot. It reasoned first. And the tokens it spends reasoning are billed at the output rate, even though you never see them.

The price tag tells the story

Google has never published parameter counts for its Flash models, so “they got bigger” is hard to prove from the outside. The pricing is easier to check.

Table 1: Gemini Flash pricing over time (per 1M tokens, standard rates)79

ModelReleasedInputOutput (incl. thinking)
Gemini 1.5 Flash2024-05$0.075$0.30
Gemini 2.5 Flash2025-06$0.30$2.50
Gemini 3 Flash—$0.50$3.00
Gemini 3.5 Flash2026-05-19$1.50$9.00
Gemini 3.6 Flash2026-07-21$1.50$7.50
Gemini 3.7 / 3.8 Flash2026-09$0.75$3.75

Input pricing rose roughly 20× between 1.5 Flash and 3.5 Flash. That 3.5 release drew its own coverage for a blunt reason: it cost three times the 3 Flash model it replaced10.

DeepSeek’s V4.1-Flash: a big model with a small name

Back to today’s release. V4.1-Flash is a mixture-of-experts model with 552B total parameters, activating 8B of them on the input side and 16B on the output side12. Its predecessor, V4-Flash, was 284B total with 13B active3.

Total parameters grew about 94% — close to a doubling. Active parameters did not follow: they became asymmetric, at 8B for input and 16B for output2. DeepSeek also attaches a separate 196B “Engram” conditional-memory module that is looked up sparsely by token match2.

The rest of the spec sheet: a 1M-token context window and up to 384K output tokens11. Thinking mode is on by default, at a default effort of “high”, with low, high, and max available12. API pricing runs $0.15–$0.30 input and $0.60–$1.20 output per million tokens, depending on peak or off-peak11.

One more detail matters. From September 14, DeepSeek will route all requests for its flagship V4-Pro to V4.1-Flash, at Flash prices1. The Flash tier is absorbing the work that used to belong to the top model.

Zhipu’s GLM-5.3-Flash: reasoning at maximum, by default

Zhipu takes the same idea further. GLM-5.3-Flash is a 320B model with 18B active, released under an MIT license with a 1M-token context window4.

The important part is the default. Its thinking mode can only be enabled, never disabled, and reasoning_effort defaults to max if you don’t set it5. Zhipu’s own guidance says to keep it at max to reproduce the published benchmark scores5.

The effect shows up in testing. Zhipu cites Artificial Analysis data showing GLM-5.3-Flash scoring 57 on the Intelligence Index v4.1.1 at $0.045 per task (discounted)4, while Artificial Analysis describes the model itself as notably slow and very verbose13. It scored 91.2 on GPQA Diamond and 84.3 on Terminal-Bench 2.1413.

Table 2: The two Flash models, side by side111245

DeepSeek V4.1-FlashGLM-5.3-Flash
Released2026-09-102026-08-26
Total parameters552B320B
Active parameters8B input / 16B output18B
Context1M1M
Thinking modeOn by default, default “high”Always on, default “max”
Input price (per 1M)$0.15–$0.30$0.15
Output price (per 1M)$0.60–$1.20$0.50
LicenseModel card publicMIT

Neither model is expensive by frontier standards. What changed is the default behavior: they no longer try to answer fast. They try to answer after thinking.

Slower by design, and the bill follows

The slowdown isn’t a quirk of one model. It follows from how reasoning models spend compute.

Epoch AI found that reasoning models use about 8× more tokens per benchmark question than non-reasoning models, and that their output lengths are growing roughly 5× per year14. A separate study found that models with nearly identical accuracy can differ by more than 26× in output length15. A model with a cheap sticker price can still produce an expensive bill.

Google now says as much in its own release notes. With Gemini 3.8 Flash, it wrote that the model “works harder,” executes extra reasoning steps on complex tasks, and “might use more tokens to maximize performance,” then advised cost-sensitive users to lower the reasoning effort or stay on 3.7 Flash16.

References

Every figure and quote in this article can be checked against the sources below:

Conclusion

“Flash” started as a useful shorthand: fast, cheap, good enough. By 2026 it has become something else — the smaller end of the frontier, where models can still be huge and where thinking is the default rather than an option.

The next time a vendor puts Flash on a model, don’t assume it’s fast. Check three numbers instead: the default reasoning effort, the active parameter count, and the cost per task. The name is marketing. Those three numbers are what you will actually feel.

Footnotes

  1. DeepSeek official news — V4.1-Flash announcement, parameter overview, V4-Pro migration https://api-docs.deepseek.com/news/news260910 ↩ ↩2 ↩3 ↩4

  2. DeepSeek-V4.1-Flash HuggingFace model card — 552B/8B/16B, 196B Engram, architecture and benchmarks https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash ↩ ↩2 ↩3 ↩4 ↩5

  3. DeepSeek-V4-Flash HuggingFace model card — previous generation at 284B total / 13B active https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash ↩ ↩2

  4. Z.ai official blog — GLM-5.3-Flash launch, 320B/18B, benchmark scores https://z.ai/blog/glm-5.3-flash ↩ ↩2 ↩3 ↩4 ↩5

  5. zai-org/GLM-5 GitHub — thinking always enabled, reasoning_effort defaults to max https://github.com/zai-org/GLM-5 ↩ ↩2 ↩3 ↩4

  6. Google official blog — Gemini 1.5 Flash announcement (“lighter-weight… fast and efficient”), May 14, 2024 https://blog.google/innovation-and-ai/products/google-gemini-update-flash-ai-assistant-io-2024/ ↩

  7. Google Gemini API pricing — Flash input/output prices by generation (output includes thinking tokens) https://ai.google.dev/gemini-api/docs/pricing ↩ ↩2 ↩3

  8. Google official blog — Gemini 2.5 thinking models, March 25, 2025 https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-model-thinking-updates-march-2025/ ↩

  9. Google official blog — Gemini 3.6 Flash at $1.50/$7.50, July 21, 2026 https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/ ↩

  10. XDA Developers — Gemini 3.5 Flash costs 3× the model it replaced, May 31, 2026 https://www.xda-developers.com/google-gemini-3-5-flash-costs-3x-model-replaced-cheap-ai-ending/ ↩

  11. DeepSeek models and pricing — context window, max output, input/output pricing https://api-docs.deepseek.com/quick_start/pricing ↩ ↩2 ↩3

  12. DeepSeek thinking mode docs — thinking on by default, default effort high, low/high/max tiers https://api-docs.deepseek.com/guides/thinking_mode ↩

  13. Artificial Analysis — independent GLM-5.3-Flash results: GPQA Diamond 91.2, described as notably slow and very verbose https://artificialanalysis.ai/models/glm-5-3-flash ↩ ↩2

  14. Epoch AI — reasoning models use ~8× more tokens; output length growing ~5× per year https://epoch.ai/data-insights/output-length ↩

  15. OckBench (arXiv preprint) — output length can differ by over 26× at similar accuracy https://arxiv.org/abs/2511.05722 ↩

  16. Google official blog — Gemini 3.8 Flash (“works harder… might use more tokens”), September 2, 2026 https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ ↩