← all stories

DeepSeek

DeepSeek's newest release is DeepSeek-V4.1-Flash, an MIT-licensed open mixture-of-experts model released on September 10, 2026, with lower off-peak API pricing. The company's earlier V4 models drew criticism in third-party agent tests, and API prices rose in August 2026.

Synthesized from 5 Yomimono stories · updated Sep 15

DeepSeek's recent output is a run of open-weight releases at falling prices. DeepSeek-V4-Flash-0731 shipped under the MIT license with weights free for commercial use, a 284 billion parameter mixture-of-experts model with 13 billion active parameters and a 1 million token context length. The same day the company notified users of a price increase for the model, and by August 6, 2026 it had raised API prices again, citing unprecedented demand. DeepSeek-V4-Flash-0731 supported commercial use from its official release, replacing a preview version.

A second model, DeepSeek V4 Pro 0813, arrived on August 14, 2026 without a formal announcement and was first noticed on social media. It has 1.6 trillion parameters, 49 billion active parameters, and a 1 million token context window, with pricing at 0.435 dollars per million input tokens and 0.87 dollars per million output tokens. The open-source coding agent Cline reported it runs at roughly one fifty-seventh the cost of the model 5 for comparable performance.

The most recent release is DeepSeek-V4.1-Flash, out on September 10, 2026. It is a mixture-of-experts system with 552 billion total parameters, 8 billion active on input and 16 billion on output, with image input support, and off-peak API pricing at half the peak rate. Third-party tests found the earlier V4 Flash model struggled on real agent tasks: Composio ran 30 difficult multi-step workflows across four harnesses and only 6 fully succeeded. That result landed after the price increases, and it is the main complication in a release slate otherwise defined by cost.

Key facts

Latest release
DeepSeek-V4.1-Flash, released September 10, 2026 as an open model under the MIT license ↗
V4.1-Flash architecture
Mixture-of-experts with 552 billion total parameters, 8 billion active on input and 16 billion on output, with image input support ↗
V4.1-Flash pricing
Off-peak API pricing is half the peak rate; KV cache tokens cut to 1/3.9 of DeepSeek-V4-Flash ↗
V4 Flash agent test results
Composio ran 30 difficult multi-step workflows across four harnesses; only 6 fully succeeded ↗
API price increase
DeepSeek raised API prices on August 6, 2026, citing unprecedented demand ↗
V4 Pro 0813 release
Released without a formal announcement, first noticed on social media; 1.6 trillion parameters, 49 billion active, 1 million token context window ↗
V4 Pro 0813 pricing
0.435 dollars per million input tokens and 0.87 dollars per million output tokens, with a discounted cache-hit input rate of 0.003625 dollars ↗
V4-Flash-0731 license and size
Open model under a commercial-use license, with 284 billion total parameters and 13 billion active parameters ↗
V4-Flash official license and cloud pricing
Weights under the MIT license; paid cloud usage listed at $0.14 per million input tokens, $0.0028 per million cached input tokens, and $0.28 per million output tokens, doubling during peak times ↗
V4-Flash context length
The official release supports a context length of 1 million tokens ↗

Timeline

Synthesized by Yomimono from the cited Yomimono stories below, each itself sourced, then editorially reviewed. Every fact links the story it came from.

Facts

Noted
licenses its model under MIT · 2026-09-11

Connections

Structured graph also available as JSON at /public/entities/deepseek. CC BY 4.0.

All coverage

Sep 11

DeepSeek Releases V4.1-Flash As An Open Model With Lower API Pricing

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026. The open model is a mixture-of-experts system with 552 billion total parameters, 8 billion active parameters on input and 16 billion on output, and it supports image input. It beat Claude Opus 5 and GPT-5.6 Sol on multiple benchmark tests. Inference efficiency cut KV cache tokens to 1/3.9 of DeepSeek-V4-Flash, and off-peak API pricing is half the peak rate.

Aug 19

DeepSeek V4 Flash Stumbles on Real Agent Tasks as Prices Surge

DeepSeek's V4 Flash model, popular for its low cost and high benchmark scores, is facing criticism after third-party tests showed it struggling on real-world agent tasks. Composio ran 30 difficult multi-step workflows across four harnesses, with only 6 fully succeeding. The company also raised API prices on August 6, 2026, citing unprecedented demand.

Aug 14

DeepSeek Quietly Releases V4 Pro 0813 Model

Chinese AI startup DeepSeek has released its latest AI model, DeepSeek V4 Pro 0813, without a formal announcement. The release was first noticed on social media and message boards, where users and tech commentators highlighted the model's performance and cost efficiency. According to the open-source AI coding agent Cline, the model scores 15.8 percent higher on Terminal Bench, a benchmark that evaluates how accurately AI can execute real system operations and development tasks on a Linux terminal, compared to the April preview model. Cline also noted that DeepSeek V4 Pro 0813 runs at roughly one fifty-seventh the cost of Claude Fable 5 while delivering comparable performance. The model has 1.6 trillion parameters, 49 billion active parameters, and a 1 million token context window. Tech writer @ChrisGPT reported benchmark improvements across Terminal Bench 2.1, CyberGym, and DeepSWE, with scores rising from 72.1 to 87.9 percent, 52.7 to 83.3 percent, and 12.8 to 62.7 percent respectively. Pricing is set at 0.435 dollars per million input tokens and 0.87 dollars per million output tokens, with a discounted cache-hit input rate of 0.003625 dollars.

Aug 9

DeepSeek Releases V4-Flash-0731 Open Model

DeepSeek has released DeepSeek-V4-Flash-0731 as an open model under a commercial-use license. The MoE model has 284 billion total parameters and 13 billion active parameters. Third-party tests by Artificial Analysis rate its intelligence on par with Google's Gemini 3.6 Flash, and it outperforms GLM-5.2 in coding and GPT-5.6 Luna in agent performance.

Aug 7

DeepSeek-V4-Flash Official Release Makes Weights Free for Commercial Use

DeepSeek has released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version. The weights are available under the MIT license, allowing free commercial use. The model is a Mixture-of-Experts type with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens. According to benchmark scores, it surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it also outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.