DeepSeek-V4.1-Flash Efficient Multimodal Model
Officially released by DeepSeek on 10 September 2026 as the smallest member of its new architecture family, with native multimodal understanding and a smaller KV cache that lowers API prices
Model Overview
Specs, access and pricing details below
Features
- ✓ New pre-training methods plus larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro
- ✓ Native multimodal visual understanding on an asymmetric architecture with a smaller KV cache, serving more users at lower cost and lowering API prices
- ✓ Available by setting the model to deepseek-flash; legacy names such as deepseek-v4-flash temporarily route to V4.1-Flash
- ✓ From 04:00 UTC on 14 September 2026 all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, until V4.1-Pro launches
- ✓ The model and the DeepSeek_V41_Tech_Report are published on Hugging Face with a stated commitment to open-source inference support; WorkBuddy (including CodeBuddy) and OpenCode support it
How to Use
- Get an API key from the DeepSeek open platform
- Set the request model to deepseek-flash (legacy deepseek-v4-flash routes over automatically)
- Submit text or image tasks to use the native multimodal capability
- To self-host, download the model and official tech report from Hugging Face, and watch for the upcoming V4.1-Pro
FAQ
What is DeepSeek-V4.1-Flash?
A model officially released by DeepSeek on 10 September 2026 as the smallest member of its new architecture family, with native multimodal visual understanding. DeepSeek says it surpasses V4-Pro across performance, cost, speed and total runtime.
How does it compare with V4-Pro?
DeepSeek says tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed and total runtime, so it plans an orderly phase-out of V4-Pro.
Do I need to change how I call the API?
Just set the model to deepseek-flash; for compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
What happens to old V4-Pro requests?
From 04:00 UTC on 14 September 2026, all deepseek-v4-pro requests route to V4.1-Flash and are billed at V4.1-Flash rates, continuing until V4.1-Pro launches.
Are the weights open?
DeepSeek says it will work with the open-source community on V4.1-Flash inference support and explore more deployment options; the model and the DeepSeek_V41_Tech_Report are published on Hugging Face.
Which official partners already support it?
DeepSeek lists WorkBuddy (including CodeBuddy) and OpenCode as official partners that fully support V4.1-Flash.