[Daily Briefing] DeepSeek 17x cheaper 💸, IBM's 100bn chip 🔬, GPT-5.6 slips 🗓️
Friday briefing. The AI price war found a new floor, IBM stacked its way past a wall the industry has been stuck on, and a major health platform is staring down an 8.8 terabyte ransom claim. Let's get into it.
DeepSeek just made its 75% price cut permanent
Here are the numbers. DeepSeek's V4-Pro now costs $0.44 per million input tokens and $0.87 per million output. GPT-5.5 charges $2.50 input and $15 output. That's more than 5x cheaper on input and roughly 17x cheaper on output, and as of this week it isn't a sale anymore. It's the price.
DeepSeek ran a 75% promo through May and everyone assumed it would lapse. It didn't. The lab has settled the model at the discount permanently, which changes the maths for anyone building on top of large language models at scale. A pipeline costing $1,500 a day on GPT-5.5 output could run nearer $90 on V4-Pro.
The honest caveat: it's a Chinese lab, so data-residency and audit rules will rule it out for plenty of enterprises, and it isn't topping every benchmark. The play most teams are landing on is a split, cheap model for the bulk and a frontier model for the few calls that truly need it.
Why it matters: A permanent price this far below the frontier sets a floor everyone else now has to argue against. The pressure lands on OpenAI: if GPT-5.6 doesn't ship a clear quality lead next month, that gap gets very hard to defend to a finance team.
Read more → (3 minute read)
Quick Hits
GPT-5.6 slipped to July. The model tipped for this week didn't land. Prediction-market odds for a 22-28 June release fell from 83% to 18%, and the IPO quiet period looks like the reason. Reported to bring a 1.5M token context window, up from 400K. Read more → (2 minute read)
One Medical hit by ransomware. ShinyHunters claims 8.8 terabytes lifted from the Amazon-owned primary care service. You can change a leaked password. You can't change your medical history. If the claim holds, it's one of the biggest health breaches of the year. Read more → (3 minute read)
IBM put 100 billion transistors on one chip. Its NanoStack design stacks transistors vertically instead of shrinking them, claiming 50% more performance for 70% less energy. That energy number is the one that makes hyperscalers move fast. Read more → (2 minute read)
OpenAI's own chip is real. Broadcom's CEO says OpenAI's Jalapeno inference chip delivers 50% lower cost per token than current Nvidia GPUs. If that holds up in production, it's the most significant internal hardware move the company has made. Big if. Read more → (2 minute read)
Prime Day's final hours. Day three is winding down and the genuinely good tech deals are thinning out fast. We pulled the ones still worth buying before it ends tonight, Sony XM6 headphones and MacBook Air M5 among them. Read more → (4 minute read)
Tool of the Day
OpenRouter
One API key that routes to hundreds of models, GPT, Claude, DeepSeek, Gemini and the open-model long tail, with live per-token pricing so you can actually compare cost per call. With DeepSeek now permanently cheap, it's the fastest way to A/B a workload across models without opening five billing accounts. Genuinely useful if you're a developer keeping your options open. Skip it if you only ever call one model, since it adds a layer and a small routing margin you don't need. Read more →
Forwarded this? Get your own free briefing at futuretechnologyhq.com/newsletter.
Genuine question for the builders reading: if DeepSeek really runs 17x cheaper on output, would you actually move a production workload onto it, or is "a lab you can't fully audit" still a hard no?
Nath, Future Technology
Some links in this newsletter may be affiliate links. We only recommend products we genuinely think are worth your time.