DeepSeek V4 Flash 0731: The Most Cost-Efficient AI Model
Released on July 31, 2026, DeepSeek V4 Flash 0731 achieves breakthrough agentic performance through re-post-training, without architecture changes. It excels in coding and agentic workflows.

Key Improvements
The 0731 update dramatically boosts benchmark scores: DeepSWE from 7.3 to 54.4 and Terminal Bench 2.1 to 82.7, outperforming much larger models. It also supports OpenAI Responses API format and reduces hallucinations.
Pricing and Efficiency
Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, with a 98% cache discount reducing cached inputs to $0.0028 per 1M tokens. This makes it up to 60% cheaper than GPT-5.6 Luna.
| Model | Input Price per 1M tokens | Output Price per 1M tokens | Cache Discount |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | $0.14 | $0.28 | 98% (cached $0.0028) |
| GPT-5.6 Luna | $0.35 | $0.70 | 90% (cached $0.035) |
Let's work together
Do you need more info, help with your project, or to develop an idea?
Whether it's an easy question, a quick doubt, or just a 5-minute chat, send me a message—it costs nothing and I'm always ready to help. I love discussing a problem to understand it, getting creative with solutions, and focusing on simple, reliable, and straightforward ideas that we can actuate quickly.
Contact me →