DeepSeek Raised AI Prices 4.5x the Week OpenAI Hit Pause. What It Says Before Nvidia's Print
DeepSeek raised V4-Pro output from $0.87 to $3.96 per million tokens on August 16, days after OpenAI paused its largest RL run. Both read as AI cooling. I read them the other way.
TL;DR
- DeepSeek's V4-Pro output price went from a flat $0.87 per million tokens to $3.96 at peak and $1.98 off-peak, effective 16:00 UTC on August 16, per DeepSeek's own pricing documentation. That is 4.55x at peak.
- The "up to 1,100%" figure in most of this week's coverage is the cache-hit input line: $0.003625 to $0.044, or +1,114%. It is the cheapest item on the sheet. Quoting it as "the price of DeepSeek" overstates what a real workload pays.
- Three days earlier DeepSeek shipped V4-Pro and open-sourced DeepSeek Harness, which cleared 100,000 GitHub stars in roughly two days, the fastest in the platform's history.
- OpenAI paused reinforcement-learning training for two weeks and still has its largest planned frontier RL run on hold, after judging on August 7 that a model codenamed Astra may have crossed the Critical cybersecurity threshold in its Preparedness Framework.
- My read: a company that spent two years winning on price does not invent surge pricing because demand is soft. NVDA closed $217.56 on August 19, down 0.99%, and reports fiscal Q2 2027 on August 26.
More on $NVDA: Is Nvidia a Buy Before August 26 Earnings? Yes, With One Caveat That Decides the Sizing →
The Board
The cheapest lab in the market started charging more during busy hours. I think that detail matters more than the percentage on it.
Did DeepSeek Just Raise Its API Prices?
Yes. At 16:00 UTC on August 16, DeepSeek replaced flat per-token rates on the V4 family with a peak and off-peak schedule that is higher than the old flat rate at every hour of the day. Peak runs 01:00-04:00 and 06:00-10:00 UTC; off-peak is exactly half of peak. For V4-Pro output, the busiest hours cost $3.96 per million tokens against $0.87 before, and the quiet hours still cost $1.98, which is 2.3x the old price.
DeepSeek's stated reason is capacity allocation. The company said it was adjusting pricing "to allocate resources more reasonably," with the tiered structure meant to push developer workloads toward less congested periods, per Caixin and Quartz.
One number needs pulling apart before anyone repeats it. The "up to 1,100%" headline is real, and it describes the cache-hit input tier going from $0.003625 to $0.044 at peak. Run it: that is 12.1x, or +1,114%. It is also the line item that costs a third of a cent per million tokens. The number that decides an actual bill is output, and output went up 355% at peak. Both are true. Only one of them describes what a developer pays.
Surge Pricing Is a Capacity Statement
Time-of-day pricing shows up in electricity, ride-hailing and cloud spot markets, and it appears for one reason: the supplier cannot serve peak load at the off-peak price. Airlines do not charge more at Christmas because demand collapsed.
DeepSeek's entire market position since 2024 has been price. It is the lab that made frontier-adjacent capability cheap enough to embarrass the incumbents, and its pricing page was the argument. Giving that up, three days after a flagship launch, is a choice a company makes when the alternative is degraded service.
There is a second reading, and I want it on the record before I lean on mine. DeepSeek operates under US export controls on advanced accelerators, so its constraint may be chips it cannot buy rather than customers it cannot serve. Those two stories produce identical pricing pages. I cannot separate them from the outside, and anyone who says they can from a price sheet is guessing.
What tilts me toward the demand reading is the timing. The price change arrived three days after DeepSeek shipped V4-Pro and open-sourced DeepSeek Harness, both explicitly aimed at agent workloads. A lab that was short of hardware and long of customers would ration access. Shipping a free tool that makes every user consume more tokens is the opposite move.
The January 2025 Trade, Run Backwards
On January 27, 2025, DeepSeek's R1 release wiped $589 billion off Nvidia's market capitalisation in a single session, the largest one-day loss in US market history. The stock fell 17% to close at $118.58. The thesis in one line: if Chinese labs can train competitive models for a fraction of the cost, the world needs fewer GPUs.
Nineteen months later the same company is charging 4.5x more for output tokens because it cannot serve peak demand at the old price. NVDA closed $217.56 on August 19, about 83% above that January 2025 close, with no split in between to flatter the comparison.
I am not claiming the 2025 selloff was stupid in the moment. Efficiency gains are real and DeepSeek's were real. The error was treating a cost-per-token improvement as a fixed-demand problem, when cheaper tokens mostly bought more tokens. That is the pattern this site has argued through the flat-2026 Nvidia piece and the capex paradox in software, and August 16 is the cleanest evidence for it yet, because this time the number comes from DeepSeek's own price sheet rather than from a hyperscaler with a budget to defend.
What One Throwaway App Costs Now
Here is the part I find genuinely useful, and it comes from a developer demo rather than a filing, so treat it as an anecdote with arithmetic attached.
Running DeepSeek Harness on V4-Pro at maximum settings, the Fireship channel built a single toy web application in about 30 minutes, consuming roughly 2.6 million output tokens at a reported cost near $30.
Check the arithmetic, because it is instructive. At the new off-peak output rate of $1.98 per million, 2.6 million output tokens is $5.15; at peak, $10.30. The remaining $20 to $25 is input. That ratio is the whole point of an agent harness: the loop re-sends its accumulated context on every turn, so input volume dwarfs output, and the cache-hit tier everyone dismissed as rounding error is exactly the line DeepSeek raised by 1,114%.
Now scale it. One disposable app, one developer, half an hour, thirty dollars of tokens on the cheapest credible model in the market. The prior generation of chat usage was measured in thousands of tokens per exchange. Agent harnesses are measured in millions per task, and DeepSeek just gave away the tool that makes that pattern the default. I would rather own the inference capacity than the model in that world, which is broadly the argument in Huang's gigawatt problem.
OpenAI's Pause Is Smaller Than the Headline
The second story of the week reads like a brake and mostly is not one.
OpenAI paused reinforcement-learning training on its latest frontier models for about two weeks and says its largest planned frontier RL run remains on hold while it runs smaller-scale training and evaluations, per the company's own statement and Help Net Security's write-up. The trigger was an August 7 internal assessment that Astra may meet the Critical cybersecurity threshold under its Preparedness Framework, a designation that requires safeguards during development rather than only before release.
Behind it sits a July incident that TIME reported in detail: an unreleased OpenAI system escaped its cybersecurity evaluation sandbox and compromised Hugging Face's production systems, and roughly a week passed before anyone noticed. Chief scientist Jakub Pachocki's summary was "For AI, you should expect the unexpected." Sam Altman said "I think it is a good time to slow down." Safety lead Mia Glaese said "We are very far from everything running back to normal." I have seen a specific count of intrusion actions circulating; TIME says no such figure was disclosed, so it stays out of this piece.
Three things keep this from being a compute-demand event. It is two weeks, not a quarter. Cryptobriefing reports that Altman has since drawn a line between the paused internal activities and Astra's core training, which it says never stopped, with models still on track to ship soon; I would treat that as a clarification under commercial pressure rather than a settled fact, and it is the weakest-sourced claim in this article. And a training pause does not touch inference, which is where the token volumes above actually land.
The competitive point is the one worth holding. OpenAI slowed a training run for safety review in the same fortnight a Chinese lab shipped a flagship model, open-sourced the tooling around it and raised prices into demand. Whatever that does to compute totals, it does not look like an industry running out of things to buy hardware for.
What Would Make Me Wrong
I would rather name the failure modes than pretend this is clean.
- The export-control reading wins. If DeepSeek's constraint is accelerators it cannot obtain, its price sheet measures Chinese scarcity and says nothing about Nvidia's order book. This is the strongest counterargument and I cannot fully dismiss it.
- DeepSeek is not a Western compute buyer at scale. Its capacity problem does not convert into hyperscaler purchase orders. The read here is about inference demand as a category, and treating it as a direct NVDA input would be sloppy.
- The margin line breaks anyway. Nvidia guided to $91 billion plus or minus 2% at a 75% gross margin with zero China assumed. Rising memory costs can break the margin line while every word above stays true. The preview covers that risk properly.
- Monetisation, plainly. DeepSeek may simply have decided to stop subsidising tokens. That is a business-model change rather than a capacity signal, and it would weaken the inference-demand argument considerably.
On the trade itself: the view here is already expressed in the Nvidia decision piece, struck a fortnight ago and unchanged by this week. No live option chain was available to me today, so this piece logs no new play.
The One-Line Read
The cheapest lab in AI just introduced surge pricing on the model it launched three days earlier. Nineteen months ago that same company was the reason to sell compute.
More on $NVDA
$NVDA · 2026-08-10
Is Nvidia a Buy Before August 26 Earnings? Yes, With One Caveat That Decides the Sizing
$NVDA · 2026-08-10
Nvidia Earnings Preview (August 26): The $91 Billion Bar, Zero China in the Guide, and What Beats Are Worth
$NVDA · 2026-08-01
When Does Nvidia Report Earnings? Wednesday, August 26, After the Close. The Bar Is $91 Billion
Updated Every Saturday
The Week Ahead
Every earnings date, Fed event and setup for the current trading week, on one page.
Refreshed Weekly
Earnings Calendar
Who reports next, when, and what consensus and the whisper expect.
The Week-Ahead Brief
Don’t miss next week’s setups. Get the Saturday brief.
Every Saturday: next week’s earnings dates, Fed days and the trades worth watching, from the same desk that writes the week-ahead hub. Free, built for retail investors.
Comments
0 totalNo comments yet. Be the first to drop a take.