Memory Is the New Bottleneck
GPUs gained 120 times more compute in a decade while memory barely moved. Positron raised $875 million betting that gap decides AI.
· 5 min read
Thomas Sohmers is the co-founder, CTO and chairman of Positron AI, which builds chips, software and full rack-scale systems for running generative AI inference. He started in semiconductors about 13 years ago, as a Thiel Fellow who founded the chip startup Rex Computing, at a time when he says silicon was "a dirty word in Silicon Valley." Positron recently raised an $875 million Series C at a $5 billion valuation from backers including Gavin Baker's Atreides Management, NEA, Valor Equity Partners and Jim Clark. Sohmers is based in Nevada and argues hard for building data centers on the empty federal desert there. These signals are from Thomas Sohmers's interview with Harry Stebbings on 20VC.
- The Memory Wall. Training is a compute problem, but inference is memory-bound: to generate every single token, the chip has to read every parameter in the model again. Between 2014 and 2024, a single Nvidia GPU gained about 120 times more compute, while memory bandwidth improved only 17 times. Part of the reason is physical. The SRAM cell on the chip has barely shrunk in about 15 years and its basic design hasn't changed in 30 or 40, so as models moved from compute-hungry CNNs to transformers, they ran straight into memory limits.
- Cached Token Margins. Processing a cached token costs roughly 1/1,000th of generating it fresh, yet providers charge a premium to write the cache and a still-profitable rate to read it. Sohmers calls the margin on those cached reads obscene and points to Anthropic's reported 80 points of gross margin on its API. He finds the "labs are burning cash" meme absurd: "If they stopped training, they'd be massively profitable overnight." He half-jokes that some of the talk about pacing the frontier is a convenient way to cut training costs ahead of an IPO.
- KV Cache Economics. SemiAnalysis traced real Claude Code sessions with dozens to hundreds of turns and found that about 96% of all tokens were cached. That turns storage into the core engineering problem. GPT-4, reportedly 1.8 trillion parameters, is about 900 GB of weights at 4-bit precision, but a single long-context user session can reach roughly 100 GB, so 50 users' sessions outweigh the model itself. Operators shuffle sessions from accelerator memory to host memory (4 to 10 times larger), then to flash and eventually slow disk, and Sohmers says the payoff is mostly cost savings for the provider rather than speed for the user.
- Context Is the Limit. Today's million-token context window holds only a fraction of Positron's own largest code repositories. Sohmers thinks the main thing stopping an agent from replacing a whole team of programmers is how much context it can hold, more than model capability. Positron's next generation will carry 8 times the memory of Nvidia's largest-memory part, while Nvidia is cutting memory per device because of market shortages. Usable context matters as much as advertised context: on a needle-in-a-haystack test, GPT-5.6 found the hidden value about 70% of the time and GPT-6 Astra over 95%.
- Compression Has a Price. The industry has dropped numeric precision from FP32 to FP16, FP8 and now FP4. Naively squeezing 16-bit values to 4 bits saves 75% of memory but can drag benchmark scores down 20% to 30%, which Sohmers calls lobotomizing the model; sharing one scale factor across groups of 16 values gets to about 4.5 bits per value with roughly 1% quality loss. Chinese labs, starved of high-memory chips by export controls, invented tricks like DeepSeek's multi-head latent attention and gated delta net, which cuts attention time by about 75%. "There's no such thing as free lunch," he says, and based on what he hears, none of the major US labs use MLA.
- Against Pacing the Frontier. Sohmers opposes the push to slow frontier AI for two reasons. Calls for a pause give ammunition to people who want to stop the technology entirely, and regulation would concentrate capability in a few hands, which he calls the modern Road to Serfdom. His deepest fear is that the big players become "the new lords and kings and everyone else is back to serfs." He thinks Dario Amodei is sincere but predicts he'll be dismayed if governments actually take over and hand control to bureaucrats who halt progress.
- China Won't Pause. A pause only works if the whole world joins, and Sohmers expects China to keep using open source to spread its models until it reaches pole position, then pull the ladder up. He thinks Western leaders are naive about staying ahead. The US military was "built to fight the last war" and would struggle against a Ukraine-style drone campaign coming across the Mexican border. He supports free trade and free exchange of ideas, with one exception: totalitarian regimes that free-ride on the open order while keeping their own markets closed.
- The Anti-Data Center Psyop. The scariest political shift to Sohmers is that opposing data centers now unites left and right, and he believes it is "almost entirely a Chinese psyop." Much of the early coverage was false: a single In-N-Out uses more water than the largest data centers in the US, golf courses use orders of magnitude more, and modern facilities run closed-loop liquid cooling. Meanwhile China adds gigawatts of mostly dirty generation and bulldozes homes to build. He would rather dress up the biggest data centers as world wonders, so that future historians see them the way we see the pyramids.
- Land and Power. More than 90% of federal land in the West is empty desert outside any national park, and Nevada has abundant geothermal and solar, yet BLM rules block data centers hundreds of miles from anyone. Sohmers says new data centers bring their own generation and could lower prices for everyone if allowed onto the grid, while power companies lobby against new capacity that would push prices down. He expects planned capacity to get built, just in different places, as blocked projects move. Alternatives exist too, such as Panthalassa's ocean-based data centers, and he never bets against Elon on space.
- Debt Over Energy. If Positron lets a buyer do in 100 MW what took 500 MW on Nvidia gear, nobody will build smaller facilities; they'll build the maximum and get more tokens per joule. "All progress is gated by energy," Sohmers says, but the binding limit is economic: how much debt the world will take on for the buildout. He worries far less about AI companies missing revenue targets than about sovereign debt and currency debasement. He trusts Oracle's business more than the US government's finances, which are propped up by the power to print currency and collect taxes.
- Local Models, More Cloud Tokens. Roughly 80% to 85% of all tokens run through the top 4 model companies and another 5% to 10% through the next 3 or 4, so Sohmers can believe about 5% ends up on premises. He argues that on-device models will increase cloud demand. A local model watching your email, calendar and messages around the clock will constantly decide to call smarter cloud models, which multiplies per-person token use well beyond what a human prompting ChatGPT occasionally generates. The next order-of-magnitude jump comes when people trust a local model to prompt the frontier models on their behalf.
- Value per Token. The Silicon Data token price index fell below $1 per million tokens this month, from about $60 five years ago, but Sohmers says a $60 token from 2021 is one "no one would pay a cent for today." Counting quality, the value per unit of intelligence has improved closer to 1,000-fold. GPT-6 Astra took an encryption block from specification through the full RTL-to-GDS chip flow on TSMC's 3nm process in a little over 50 hours, running above 1 GHz, work he would expect to take a new engineer 2 to 3 weeks. He expects per-token pricing to stick because margins are easy to calculate, though labs may eventually sell an unlimited virtual worker for a flat fee, something like $1 million a year.
Watch the full video at https://www.youtube.com/watch?v=6ohZuFkq-aU.