DeepSeek Put V4 Flash in Public Beta. Cheap Agent Scores Still Need Your Evals.

July 31 public beta of V4-Flash-0731: vendor agent benches beat Pro preview at roughly one-third the output price—while Nikkei flags a heating price war and peak-hour pricing plans.

calender-image
August 1, 2026
clock-image
6 min read
DeepSeek Put V4 Flash in Public Beta. Cheap Agent Scores Still Need Your Evals.
Free weekly briefingThe Business AI Briefing for people who run the Business — 5 min, zero hype.
Get the briefing free →

Via MarkTechPost: DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains · Also: Nikkei Asia on the V4 beta and AI price war

Public beta, cheaper loops, louder price war

On July 31, 2026, DeepSeek moved the official DeepSeek-V4-Flash-0731 checkpoint into the open and put the deepseek-v4-flash API into public beta. MarkTechPost's technical write-up is clear: architecture and size are unchanged from the preview—the jump is re-post-training. Same day, Nikkei Asia framed the release as fuel for a global AI model price war, noting DeepSeek's rates sit far below many U.S. rivals and that the company says it will implement a peak-hour pricing plan.

That is the SME story in one invoice. Vendor-reported agentic scores jumped—Terminal Bench 2.1 to 82.7 from 61.8; DeepSWE to 54.4 from 7.3; Toolathlon-Verified to 70.3—beating the V4-Pro preview prints DeepSeek published, at roughly one-third Pro's output price ($0.28 vs $0.87 per million output tokens). Flash input sits at $0.14 per million on a cache miss and $0.0028 on a hit, with a 2,500 concurrency limit versus 500 for Pro.

None of that is a permission slip to point every workflow at Flash and declare victory. Agent scores are harness-sensitive. MarkTechPost notes Code Agent tasks used an unreleased DeepSeek Harness in minimal mode, and some DSBench sets are internal. Cheap tokens that fail your documents are not cheap.

Why the Nikkei price-war frame matters to a 40-person firm

Owner-led firms do not need another Silicon Valley panic narrative. They need a routing rule that survives a Friday release note. Nikkei's Hong Kong report puts the beta in a competitive market story: more capable offerings, aggressive pricing, and a planned peak-hour schedule. MarkTechPost supplies the operator detail—Responses API support on Flash, Codex adaptation, unchanged V4-Pro / app / web surfaces for this drop.

Put those together and you get three business facts:

  • The menu moved under the same model ID. Callers already on deepseek-v4-flash get the 0731 upgrade without renaming—continuity is easier, regression testing is still required.
  • Flash is priced like a fleet engine. High concurrency and low output rates invite agent loops. That is useful for research, drafting, and tool-using workflows that clear your quality bar.
  • Peak-hour pricing is a forecast risk, not a folklore rumor. Nikkei reports DeepSeek says it will implement peak-hour pricing. Treat the effective date and multipliers as something to confirm on DeepSeek's pricing docs before you lock a budget model. Flat rates today are not a promise about next quarter's Tuesday afternoon in Beijing time.

Self-hosting remains a different product. Weights are MIT-licensed and ungated, but Unsloth's 3-bit path still needs roughly 110 GB combined RAM and VRAM. Most SMEs will buy API minutes, not a GB300 node. That is fine—as long as someone owns the fallback when the cheap endpoint changes validation, rate limits, or the clock on peak windows.

Blog Image

What smart firms do with a hotter mid-tier

Smart firms treat Flash-0731 as a procurement input, not a loyalty program.

  1. Re-run your top ten agent and coding tickets. Vendor benches are a headline. Your documents are the exam. Score rework time, tool-call reliability, and citation behavior—not vibes from a demo reel.
  2. Price the finished task, including peak windows. Model $0.28/M output as the base case. Add a scenario where peak-hour pricing lands. If your load sits in Beijing business hours, that scenario belongs in the spreadsheet now.
  3. Keep Pro—or another provider—as a named fallback. Flash may win volume work. High-stakes judgment, regulated drafting, and anything that failed your Flash bake-off stay on a higher tier until evidence says otherwise.
  4. Separate API convenience from self-host romance. MIT weights are nice. A 110 GB memory floor is still a capital decision. Most owner-led firms should not confuse open weights with free lunch.
  5. Assign an owner for DeepSeek release notes. Same model ID, new checkpoint, new API formats, future peak clocks—someone has to read the page so the partner does not discover breakage in a client chat.

"A cheaper Flash model posting better vendor agent scores than Pro is a routing memo—not a reason to skip your evals or ignore peak-hour pricing plans." — AgentsROI on DeepSeek V4-Flash-0731

How AgentsROI turns a price war into a routing policy

AgentsROI.ai is a managed AI services provider for owner-led SMEs. We do not sell DeepSeek. We help you decide when Flash is the right seat—and what happens when the price board or the model card moves again.

Model Selection & Continuity Planning is the primary fit. We map jobs to tiers using your workloads, set a default and a fallback, and write decision criteria so the next re-post-training drop does not restart the argument from zero. When Artificial Analysis, MarkTechPost, or Nikkei publish the next print, you already know which workflows are allowed to move.

Managed AI Operations keeps the cheap default honest: usage watch, peak-hour scenario updates, smoke tests after alias or checkpoint changes, and permission reviews when agent capabilities jump. Vendor-neutral by design—open weights, commercial APIs, or a mix. The goal is finished work at acceptable quality and a cost curve you can explain without theology about which lab is “winning.”

Buy the loop, watch the clock

DeepSeek's July 31 public beta of V4-Flash-0731—stronger vendor agent scores at roughly one-third Pro output pricing, framed by Nikkei as part of a heating price war with a peak-hour plan ahead—is a routing memo for owner-led firms. Use it where your evals clear. Keep a fallback. Confirm peak-hour effective dates before you forecast next quarter like the rate card never changes.

If you want a Model Selection pass on your agent and coding stack, book a no-pressure assessment.

This article summarizes publicly reported information from MarkTechPost and Nikkei Asia and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change—including announced peak-hour pricing plans whose effective dates may still be pending. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.