DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now costs. The model, tagged 0813, went live on OpenRouter this week and climbed past 540 points on Hacker News within a day.
The headline claim is that V4 Pro reaches 87.9 on Terminal-Bench 2.1, placing it a tenth of a point behind Fable 5 on that evaluation. It does so at $0.435 per million input tokens and $0.87 per million output tokens, a fraction of what comparable frontier models charge.
The rollout was quick and largely unannounced. The model surfaced on aggregators on August 12, ahead of DeepSeek's own launch late the following evening Beijing time, a sequencing that put hands-on results in front of developers before any formal announcement circulated.
Architecture and Limits
V4 Pro is a mixture-of-experts model with roughly 1.6 trillion total parameters, of which about 49 billion are active during inference. That ratio is the core of the cost story: the model carries frontier-scale knowledge while routing each token through a small fraction of its weights.
The context window runs to 1 million tokens with a maximum output of 384,000 tokens. DeepSeek also prices cache hits separately at roughly $0.0036 per million input tokens, which matters substantially for agentic loops that resend a large, mostly static context on every turn.
OpenRouter lists the model with tool calling and structured output support, and routes requests across multiple hosting providers using selectable modes that optimize for balanced cost and speed, raw throughput, or tool-calling accuracy. That multi-provider arrangement gives the model redundancy that single-vendor endpoints lack.
Large Gains Over the Preview
The jump from the preview build to general availability is where the numbers get unusual. On DeepSWE, the score moved from 12.8 to 62.7. CyberGym rose from 52.7 to 83.3. Terminal-Bench 2.1 climbed from 72.1 to 87.9.
Improvements of that magnitude between a preview and a general release are uncommon, and suggest the preview was an early checkpoint rather than a near-final candidate. The pattern concentrates in agentic and tool-use evaluations rather than static knowledge tests, which aligns with where the broader industry has focused its post-training effort this year.
The Cost Argument
Positioned against Western frontier models, the pricing gap is the point. At $0.435 and $0.87 per million tokens, V4 Pro sits well below the tiers most labs reserve for their strongest models, and roughly in line with what competitors charge for mid-sized offerings.
Independent trackers frame the gap even more sharply, with at least one comparison pegging V4 Pro at a small fraction of the per-token cost of the model it narrowly trails on Terminal-Bench. Several hosts, Together AI among them, now serve the model alongside DeepSeek's own endpoint.
For teams running agents at volume, where a single task may consume hundreds of thousands of tokens across many turns, that difference determines which workloads are economically viable at all. A benchmark deficit of a tenth of a point matters far less than an order-of-magnitude difference in unit cost.
What to Watch
Benchmark scores and production behavior are not the same thing, and vendor-reported evaluation figures warrant independent replication before teams commit. Throughput, latency under load, and tool-calling reliability across providers will determine whether V4 Pro holds up outside controlled tests.
Still, the release continues a pattern DeepSeek has established repeatedly: ship a model that lands within noise of the frontier, then price it far enough below the market to force a response. Availability through OpenRouter and other aggregators means developers can route traffic to it without changing much beyond a model identifier.






