The Leaderboard Is Not the Territory
Part 2 of a series on AI benchmarking: what happens when a measurement becomes a target, and why a team of Berkeley…
Part 2 of a series on AI benchmarking: what happens when a measurement becomes a target, and why a team of Berkeley…
Part 1 of a series on AI benchmarking: why a 2-point gap on a leaderboard tells you almost nothing, and how the…
A Shanghai lab just shipped a model that does at 1M tokens what full attention can’t do at any price. The architecture…
Until now, training a 100B+ parameter model required a cluster, a multi-million dollar budget, and a very patient CFO. A new paper…
Three papers, three institutions, one uncomfortable conclusion for the AI industry There is a prevailing assumption baked into how we talk about…
NVIDIA’s RTX Spark small desktop is a bet that the future of AI agents isn’t in the cloud. It’s plugged into the…
Researchers just proved that LLMs can deanonymize pseudonymous users at scale with off-the-shelf tools and a sandwich budget. Here’s what that actually…
A deep dive into the AI strategy buried inside SpaceX’s S-1 filing: data centers, frontier models, space-based compute, chip factories, and an…
A breakdown of the Q1 2026 global LLM market: who’s winning, who’s bluffing, and why user counts are the most misleading number…
Arthur Mensch’s landmark hearing before the National Assembly laid out a stark vision and a harder question: does Europe have the will…