2026-07-31 · 1 min read

DeepSeek V4 Flash 0731: The Most Cost-Efficient AI Model

Released on July 31, 2026, DeepSeek V4 Flash 0731 achieves breakthrough agentic performance through re-post-training, without architecture changes. It excels in coding and agentic workflows.

article image

Key Improvements

The 0731 update dramatically boosts benchmark scores: DeepSWE from 7.3 to 54.4 and Terminal Bench 2.1 to 82.7, outperforming much larger models. It also supports OpenAI Responses API format and reduces hallucinations.

Pricing and Efficiency

Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, with a 98% cache discount reducing cached inputs to $0.0028 per 1M tokens. This makes it up to 60% cheaper than GPT-5.6 Luna.

Pricing Comparison
ModelInput Price per 1M tokensOutput Price per 1M tokensCache Discount
DeepSeek V4 Flash 0731$0.14$0.2898% (cached $0.0028)
GPT-5.6 Luna$0.35$0.7090% (cached $0.035)
← Scroll right to see more →

Let's work together

Do you need more info, help with your project, or to develop an idea?

Whether it's an easy question, a quick doubt, or just a 5-minute chat, send me a message—it costs nothing and I'm always ready to help. I love discussing a problem to understand it, getting creative with solutions, and focusing on simple, reliable, and straightforward ideas that we can actuate quickly.

Contact me

Switch Topic

Choose a specialized topic to explore: