DeepSeek V4.1 Flash Introduces Native Vision With Reduced API Pricing

Published:

DeepSeek has officially introduced V4.1 Flash, replacing its previous V4 Flash and V4 Flash Vision Experimental models while introducing lower API pricing and built-in image understanding capabilities. The new model became available on September 10, 2026, under the API name deepseek-flash, with existing Flash model names continuing to function as aliases that automatically route requests to the latest version. The launch combines improvements in performance, multimodal processing, and cost efficiency, providing developers with a single model designed for both text and image workloads while simplifying migration from earlier Flash releases.

DeepSeek V4.1 Flash is built on a redesigned Mixture of Experts architecture featuring approximately 748 billion total parameters, including a 552 billion parameter backbone and 196 billion Engram parameters. Despite its overall size, the model activates only around 8 billion parameters during prefill and approximately 16 billion during decoding, helping reduce inference costs without sacrificing performance. The model introduces native multimodal support through the DeepSeek ViT encoder, allowing it to understand both images and text with support for image resolutions of up to approximately 1344 × 1344 pixels. It also offers a context window of up to one million tokens and can generate outputs of up to 384,000 tokens. Developers can further customize performance by adjusting reasoning effort on a scale from 1 to 100, enabling them to balance response quality, speed, and operating costs according to their application requirements.

According to DeepSeek, V4.1 Flash delivers notable improvements across coding, automation, and AI agent benchmarks compared to the previous V4 Flash model and, in several areas, also performs competitively against V4 Pro. The company reported that its DeepSWE score increased from 54.4 to 74.2, while Terminal Bench 2.1 improved from 82.7 to 90.6. Additional benchmark results include scores of 88.1 on CyberGym and 54.8 on AutomationBench. Independent evaluations have also produced encouraging results, with Artificial Analysis reporting a 68.9 percent score on AutomationBench AA and Vals.ai assigning the model a Vals Index score of 57.86, placing it among the leading open weight models in that comparison. However, DeepSeek noted that V4.1 Flash does not lead every benchmark, as some frontier AI models continue to perform better on selected Terminal Bench evaluations as well as specialized tests such as SRE Bench and Harvey Legal Agent Benchmark.

The release also introduces significant reductions in API pricing. During off peak hours, cached input is priced at $0.003 per million tokens, while uncached input costs $0.15 per million tokens and output is priced at $0.60 per million tokens. Peak usage pricing doubles those rates to $0.006 for cached input, $0.30 for uncached input, and $1.20 for output. The largest reduction applies to cached input, making the model more economical for AI agents and applications that repeatedly reuse large context windows. DeepSeek also announced that beginning September 14, 2026, requests sent to deepseek-v4-pro will automatically be redirected to V4.1 Flash and billed using the new Flash pricing until V4.1 Pro becomes available. While the transition is expected to reduce costs for existing V4 Pro users, DeepSeek noted that the updated architecture may produce differences in prompting behavior, tool calling, and response style compared to the previous model.

Follow the SPIN IDG WhatsApp Channel for updates across the Smart Pakistan Insights Network covering all of Pakistan’s technology ecosystem.

Related articles

spot_img