DeepSeek-V4.1-Flash

Last updated 2026-09-25. V4.1-Flash (Sep 2026)

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone, a Causal Encoder-Decoder that activates 8B parameters per token during prefill and 16B during decode, a 1,048,576-token context, Engram conditional memory (196B parameters), and a global KV cache of 890 bytes per token using FP4 (E2M1) caching. API model name: deepseek-flash.

Parameters
552B Backbone; active 8B Prefill / 16B Decode
Architecture
Asymmetric Causal MoE (6 of 384 experts) + Engram (196B) + FP4 E2M1 KV Cache
Context
1,048,576 Tokens (1M Native)
KV cache
890 bytes / token (FP4 E2M1)
Peak input / 1M (cache miss)
0.3
Peak output / 1M
1.2

Price source: DeepSeek Models & Pricing. Peak cache miss $0.30 / off-peak $0.15 per 1M input. Peak cache hit $0.006 / off-peak $0.003. Peak $1.20 / off-peak $0.60 per 1M output.

Sourced benchmarks

  • Codeforces (Rating): 3,471 — Instruct-model Codeforces rating at reasoning_effort=100, temperature 1.0, top_p 0.95. Source
  • GPQA Diamond (Pass@1): 90.9% — GPQA Diamond pass@1 at maximum reasoning effort. Source
  • Terminal-Bench 2.1 (Pass@1): 90.6% — Code-agent benchmark using the Minimal mode of DeepSeek Harness. Source
  • HLE with tools (Pass@1): 63.9% — Humanity's Last Exam with tools, not the text-only number. Source

All models · API pricing · FAQ