← Back to AI Tools

DeepSeek-V4.1-Flash Efficient Multimodal Model

Officially released by DeepSeek on 10 September 2026 as the smallest member of its new architecture family, with native multimodal understanding and a smaller KV cache that lowers API prices

Model Overview

Specs, access and pricing details below

Features

  • ✓ New pre-training methods plus larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro
  • ✓ Native multimodal visual understanding on an asymmetric architecture with a smaller KV cache, serving more users at lower cost and lowering API prices
  • ✓ Available by setting the model to deepseek-flash; legacy names such as deepseek-v4-flash temporarily route to V4.1-Flash
  • ✓ From 04:00 UTC on 14 September 2026 all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, until V4.1-Pro launches
  • ✓ The model and the DeepSeek_V41_Tech_Report are published on Hugging Face with a stated commitment to open-source inference support; WorkBuddy (including CodeBuddy) and OpenCode support it

How to Use

  1. Get an API key from the DeepSeek open platform
  2. Set the request model to deepseek-flash (legacy deepseek-v4-flash routes over automatically)
  3. Submit text or image tasks to use the native multimodal capability
  4. To self-host, download the model and official tech report from Hugging Face, and watch for the upcoming V4.1-Pro

FAQ

What is DeepSeek-V4.1-Flash?

A model officially released by DeepSeek on 10 September 2026 as the smallest member of its new architecture family, with native multimodal visual understanding. DeepSeek says it surpasses V4-Pro across performance, cost, speed and total runtime.

How does it compare with V4-Pro?

DeepSeek says tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed and total runtime, so it plans an orderly phase-out of V4-Pro.

Do I need to change how I call the API?

Just set the model to deepseek-flash; for compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.

What happens to old V4-Pro requests?

From 04:00 UTC on 14 September 2026, all deepseek-v4-pro requests route to V4.1-Flash and are billed at V4.1-Flash rates, continuing until V4.1-Pro launches.

Are the weights open?

DeepSeek says it will work with the open-source community on V4.1-Flash inference support and explore more deployment options; the model and the DeepSeek_V41_Tech_Report are published on Hugging Face.

Which official partners already support it?

DeepSeek lists WorkBuddy (including CodeBuddy) and OpenCode as official partners that fully support V4.1-Flash.