Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 2 sources · 1h ago ·

DeepSeek-V4-Flash Official Release Makes Weights Free for Commercial Use

The release came hours after OpenAI cut GPT-5.6 Luna prices by 80 percent, and DeepSeek has now notified users of a significant price increase for the model, suggesting the low-cost strategy drew demand that strained its computing resources.

Key Facts

  • DeepSeek released DeepSeek-V4-Flash-0731 as an official open model under a commercially usable license on July 31, 2026.
  • The model has 284 billion total parameters, activates 13 billion during inference, and supports a 1 million token context length.
  • Unsloth released GGUF-format quantized versions of DeepSeek-V4-Flash in 13 steps from 1-bit at 82.5GB to 8-bit at 162GB.
  • DeepSeek's API pricing for deepseek-v4-flash is $0.14 per million input tokens, $0.0028 per million cached input tokens, and $0.28 per million output tokens, doubling during peak times.
  • DeepSeek notified users on August 6, 2026 of a significant price increase for DeepSeek-V4-Flash-0731 without specifying new amounts.

Reporting from 2 sources: GIGAZINE, GameBusiness.jp.

DeepSeek-V4-Flash Official Release Makes Weights Free for Commercial Use

DeepSeek has released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version. The weights are available under the MIT license, allowing free commercial use. The model is a Mixture-of-Experts type with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens. According to benchmark scores, it surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it also outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.

  • DeepSeek notified users of the price increase on August 6, 2026, without specifying the new amount, and asked them to plan usage accordingly.
  • The price hike follows reports of the model's inference speed occasionally slowing to extreme levels as demand surged, with DeepSeek's API status log showing frequent access difficulties around the release.
  • CEO Liang Wenfeng said in July 2026 that DeepSeek's computing resources equal about 20,000 NVIDIA H100 GPUs, which some reports say the demand spike has strained.
  • Bloomberg reports DeepSeek is in the middle of a large fundraising round, targeting about $8 billion, with a post-round valuation near $74 billion, and a possible IPO within 2026.
  • The fundraising is the second round, resumed on August 6 after trading was temporarily halted in late July following the leak of the CEO's investor remarks on US-China AI competition.
  • OpenAI cut GPT-5.6 Luna prices by 80 percent on July 31, bringing input to $0.20 per million tokens and output to $1.20 per million tokens, hours before DeepSeek's release.
  • OpenAI said the cut came from extracting further architectural efficiency, though some observers called it a pricing strategy aimed at Chinese AI labs.
  • The price-cut GPT-5.6 Luna became more cost-efficient than closed models like Claude Sonnet 5 and Gemini 3.6 Flash, as well as Chinese open models including DeepSeek V4 Pro and GLM-5.2.

Synthesized by Yomimono from the 2 cited sources below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources