Tech news & tech videos

LeepCast

The signal in the noise. Daily tech news, deep dives, and original videos on AI, software, and the future of building.

Latest tech news

The biggest stories from the last 24 hours.

AIJun 27, 2026

OpenAI sunsets GPT-4.5 and auto-migrates ChatGPT users to GPT-5.5

On June 27 OpenAI finished a 30-day sunset it had flagged in late-May release notes, pulling GPT-4.5 from ChatGPT and dropping it out of the model picker. Existing conversations move automatically to the matching GPT-5.5 model, so users don't have to lift a finger. It's a routine housekeeping event, but a pointed one: a model that felt frontier not long ago is now retired to clear room for a newer tier.

Why it matters: The consumer side of this is a non-event — chats migrate silently — but the lesson for builders is the opposite of frictionless. If you ship on these APIs, a model version is not a stable foundation; it's a depreciating asset with a sunset clock you don't control. Thirty-day windows and forced migrations mean version pinning, regression-testing against successors, and a standing migration playbook are now table stakes, not optional hygiene. The deeper trend is that model lifecycles are compressing faster than most product roadmaps, so prompts and eval suites tuned to one model's quirks become liabilities the moment it's deprecated. My take: teams that treat the model as a swappable dependency behind an abstraction layer will absorb these transitions calmly, while teams that hard-code behavior to a specific version will keep getting surprised.

OpenAI
AIJun 26, 2026

OpenAI previews GPT-5.6 as a three-model family — Sol, Terra, and Luna

OpenAI's GPT-5.6 preview isn't one model but a lineup: Sol as the flagship, Terra as the balanced everyday option, and Luna as the fast, cheap one. What stands out is less the headline capability claims — stronger agentic coding, biology, and cybersecurity, a new max reasoning effort, and an ultra mode that spawns subagents — and more the release choreography. OpenAI is handing the models first to roughly 20 organizations in a staged rollout, after briefing the US government on the models and its plans, with wider access promised in the coming weeks.

Why it matters: The tiered lineup formalizes something builders have been improvising for a year: you don't route every request to the biggest model, you match the model to the job and the budget. Publishing that split as a product decision means capability-per-dollar becomes a first-class design axis, and 'which tier' becomes a routing question you have to answer in your own stack. The more consequential change is the access model. A frontier launch that goes to ~20 orgs after a government briefing tells you the newest capabilities now arrive behind a staged, policy-shaped gate rather than a public API on day one. If you're building on the frontier, plan for a world where the strongest models are a privilege you wait for, not a switch you flip — and where your competitors' access to them may be uneven for weeks at a time.

OpenAI
ChipsJun 25, 2026

Qualcomm unveils the Dragonfly C1000, a data-center CPU built for agentic AI orchestration

At its Investor Day, Qualcomm introduced the Dragonfly C1000 — a 250-plus-core, 5GHz data-center CPU aimed not at training but at the orchestration side of agentic AI: the high-throughput sequential reasoning and constant context-switching that GPUs handle poorly and fast CPUs handle well. The chiplet design supports PCIe Gen7 and CXL and claims roughly 2x better performance per watt than incumbents. Mark Zuckerberg confirmed Meta signed a multi-generational supply agreement, with production slated for the second half of 2028; Qualcomm also said it's acquiring AI-software firm Modular and deepening ties with Hugging Face.

Why it matters: The pitch reflects a real architectural shift: as workloads move from one-shot chatbot calls to long-running agents that plan, branch, call tools, and juggle context, the bottleneck migrates from raw matrix multiplication to coordination — and coordination is a CPU-shaped problem, not a GPU one. Building a chip explicitly for that is a direct shot at Intel, AMD, and the GPU-centric assumption that everything in the AI data center is a matmul. Landing Meta as a launch customer is the part that makes it credible rather than aspirational; hyperscaler commitment is what turns a spec sheet into a market, and it signals the agentic data center is becoming its own hardware category with its own procurement logic. The obvious caveat is time — second-half-2028 production is far enough out that the workload it's tuned for could look different by the time it ships, and the CPU-vs-GPU division of labor in agent serving is still being figured out. Read alongside the Modular acquisition and the Hugging Face tie-up, though, the strategy is coherent: Qualcomm is assembling a full software-plus-silicon stack for a post-GPU-monoculture data center.

CNBC
AIJun 25, 2026

Anthropic accuses Alibaba's Qwen lab of the largest known "distillation attack" on Claude

A letter Anthropic sent the US Senate Banking Committee, made public this week, accuses Alibaba's Qwen lab of what it calls the largest known distillation attack against Claude — roughly 25,000 fraudulent accounts and about 28.8 million interactions between April 22 and June 5, allegedly to extract Claude's most advanced software-engineering and agentic-reasoning behavior. Distillation here means systematically prompting a rival model and training on its outputs to replicate its behavior, which Anthropic says breached its terms. Senators reacted by floating an amendment to sanction foreign firms that improperly access US model outputs.

Why it matters: Distillation is a normal, widely used technique — the entire premise of many smaller open models is learning from a stronger one's outputs — so the fight isn't about the method, it's about doing it to a competitor at industrial scale and against its terms. That reframes model behavior as intellectual property that can allegedly be stolen through the front door of an API, which raises a genuinely hard question the industry hasn't settled: if capabilities can be siphoned by anyone with enough accounts and prompts, what exactly does a frontier lead protect, and for how long? The move to Washington is the real escalation — routing this through a Senate committee and a sanctions amendment folds model-output IP into US-China tech tensions and invites regulation of what labs are permitted to learn from each other. Expect a defensive arms race in response: tighter rate limits, aggressive anti-abuse detection, and stricter terms of service, all of which quietly raise the cost and friction of legitimate API use for ordinary builders caught in the blast radius.

CNBC
ChipsJun 24, 2026

Qualcomm to acquire Modular for ~$3.9B in a bid to make Nvidia's CUDA moat optional

Qualcomm's reported $3.9 billion move for Modular is a software play wearing a hardware company's clothes. The asset is Modular's Mojo language and MAX inference engine, which let developers write inference code once and run it optimized across Nvidia, AMD, Intel, Qualcomm, and Apple silicon — a direct swipe at the CUDA lock-in that keeps AI workloads tethered to Nvidia. The deal brings roughly 150 people, including co-founders Chris Lattner (of LLVM and Swift) and Tim Davis, and is expected to close in the second half of 2026.

Why it matters: Nvidia's durable advantage was never only faster chips; it's that virtually all AI runs on CUDA, so switching hardware means rewriting your stack. A portability layer attacks that at the root — if 'write once, run anywhere' actually holds up in production, the cost of leaving Nvidia drops and hardware becomes a per-workload price comparison instead of a lifetime commitment. That's why the buyer is a chipmaker: Qualcomm doesn't need to out-engineer Nvidia's GPUs if it can commoditize the software that locks customers in. The broader signal is that the AI-chip war is climbing the stack, from silicon to the compiler and runtime that decide what silicon you're even allowed to consider. The open question is adoption — abstraction layers have promised hardware freedom before and mostly delivered a leaky lowest common denominator; the founders' pedigree is the reason to take this attempt seriously.

CNBC
ChipsJun 24, 2026

OpenAI and Broadcom tape out Jalapeño, a custom inference chip in about nine months

OpenAI and Broadcom pulled the cover off Jalapeño, OpenAI's first custom silicon — a reticle-sized ASIC built to run large models rather than train them. The companies say it went from design to manufacturing tape-out in roughly nine months, one of the fastest advanced-ASIC cycles they've seen, with OpenAI's own models pitching in on the design. Engineering samples are already running lab workloads including a GPT-5.3 Codex variant, with performance-per-watt claimed well above current state of the art, as the first step in a multi-generation OpenAI-Broadcom-Celestica platform slated to start deploying by the end of 2026.

Why it matters: Designing the chip your models run on is the clearest statement yet that frontier labs no longer want to rent their entire stack. Inference — not training — is where cost compounds as usage scales, so a purpose-built inference ASIC with better performance-per-watt attacks the single biggest line item in serving models at OpenAI's volume. The nine-month cycle is the eyebrow-raiser: if using their own models to compress ASIC design timelines holds up, it hints at a flywheel where AI accelerates the hardware that makes AI cheaper. Strategically this is another pressure point on Nvidia's margins and a hedge against GPU allocation being the thing that caps a lab's growth. The caveats are real — engineering samples aren't volume production, and end-of-2026 deployment leaves ample room for slippage — but the direction of travel is unmistakable: owning silicon is becoming a frontier-lab requirement, not a luxury.

Tom’s Hardware

The LeepCast brief

One email a week. The stories and videos that actually matter, no filler. Join thousands of builders who stay ahead.

No spam. Unsubscribe anytime.