Open Weights, Local Silicon, Self-Governance
Kimi K3 puts near-frontier open weights in builders' hands, Apple's rumored M7 Ultra points at local trillion-parameter inference, and Hassabis's AI SRO is how the private sector keeps pace.
Three things landed in the same news cycle this week. A near-frontier open-weights model. A rumor that Apple Silicon will soon hold enough unified memory to run trillion-parameter models locally. And a concrete proposal for how the industry should police the frontier. They sound like different stories. They are the same one: capability is leaving the closed-lab clubhouse, and the question is whether builders and the industry catch up — technically and institutionally — before the gap closes the hard way.
Kimi K3
Moonshot AI shipped Kimi K3 mid-week: a 2.8 trillion parameter Mixture-of-Experts model that activates a small slice of experts per token (on the order of tens of billions active), with a 1M-token context window and native multimodal inputs. On Arena frontend coding it jumped to the top of the board; on broader text composites it sits just behind the closed frontier — Claude Fable 5 and GPT-5.6 Sol — which is exactly the framing Moonshot itself uses. Not “we won AI.” More honest: we closed most of the gap, and we are going to publish the weights.
That last part is the strategic fact — and it has not happened yet. The API is live now (kimi-k3 on Moonshot’s OpenAI-compatible endpoint). Downloadable weights are pledged for the end of July, not released. Until the checkpoint is public, treat “open source” headlines carefully. What matters for builders is open weights: the right to run, inspect, fine-tune, and host without renting the lab’s inference stack forever. Full 2.8T self-host is cluster-scale, not a laptop. The permission still changes the game. Agents, long-horizon coding, air-gapped enterprise, custom fine-tunes — the surface area expands the moment the weights land.
David Sacks’s take is blunt: if the U.S. ties itself in knots on regulation and data centers while open Chinese models take leaderboards, that is how you lose the race. The builder version of the same point is simpler. Permissionless models compress the closed-lab advantage. Pricing pressure follows. So does the incentive to ship agents on top of weights you control.
Apple M7 Ultra
Meanwhile, the rumor mill — via Bloomberg’s Mark Gurman and the usual secondary coverage — says Apple is designing an M7 Ultra for around 2028, aimed at Mac Studio-class machines, with up to 1.5TB of unified memory. Whether Apple ships the full config depends on the memory market. Treat it as a rumor until silicon ships.
I feel the bottleneck daily. My Mac Mini runs my local models and is already choking on memory — models that do not fit, swap thrash, inference that stalls when the context gets long. Memory is the constraint. At 8-bit quantization, a trillion-parameter dense model wants on the order of a terabyte-plus of RAM; at 4-bit, roughly half that. Apple’s unified memory architecture has already made large open models practical on Mac Studios today. Doubling toward 1.5TB is not “Apple trains GPT.” It is Apple enabling the open-weights stack — privacy, offline, no per-token cloud tax — on hardware builders already own.
Put next to Kimi K3, the timeline gets interesting. Open MoE giants will still want clusters for full fidelity. But quantized, distilled, and expert-routed variants — the models builders actually run — get more viable on a single desk every generation. Local is not a hobby niche. It is becoming the default for anyone who cares about IP, latency, or cost at the edge.
Hassabis and an AI SRO
Which brings us to governance. Demis Hassabis published A Framework for Frontier AI and the Dawning of a New Age (July 14, 2026) — a call for a U.S.-led, industry-funded standards body modeled on FINRA. The core of the proposal, in his words:
The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organisation, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives. Funding would need to be substantial and likely mostly come from industry, in order to attract world-class technical talent and provide the necessary compute resources for large-scale testing.
Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow, meaning that Frontier Models would be required to pass it to be deployed in the US market.
The framework could apply to Frontier-class models no matter their country of origin or whether they are open or closed, but any non-frontier models, say from startups or academia, would be exempt from this process.
That is the right shape: independent technical experts, open-source representation, government oversight. Cyber, bio, and deception testing before release — voluntary first, then required for U.S. deployment once the protocol works. Non-frontier stays out of scope. Open and closed both count if they clear the capability bar.
I am supportive. Government cannot do this alone. It cannot hire the talent, buy the compute, or update benchmarks on frontier cadence. The private sector has to govern itself — with real technical horsepower and public accountability — or we get either theater regulation that slows builders, or no guardrails at all at the edge that matters.
Here is the weave. Open weights like Kimi K3 put near-frontier capability in more hands. Silicon like the rumored M7 Ultra puts more of that capability on machines outside the lab. That is good for builders. It is also why a competent AI SRO is not optional theater. You want safety evaluations that keep up with the models — and a bright line that leaves everyone below the frontier free to ship. Capability is dispersing. The industry has to set the rules of the road.
Husband, Father, Friend, Technologist, Entrepreneur and Amateur Humorist
Continue reading

AI Sovereignty: Can we trust Sam and Dario?
The AI wave is real and AGI may already be here — but enterprise trust in frontier labs is fracturing. A builder's tour of Sequoia's computation revolution, Karp's sovereignty warning, and the open-source path forward.


Adios WordPress
After almost twenty years on WordPress.com, Memory Leak moves to a site built for SEO, AEO, and GEO — with agents as first-class citizens and an agentic development lifecycle behind the rebuild.