The Parallel Model Wins at Inference
Today's chatbots write one word at a time. Stanford's Stefano Ermon is betting the next generation writes them all at once.
· 4 min read
Stefano Ermon is a Stanford computer science professor and the co-founder and CEO of Inception, which builds diffusion-based language models under the Mercury name. His lab produced score-based generative models in 2019, the work that became the basis of modern image diffusion and fed into Stable Diffusion and Midjourney. He co-advised the FlashAttention work, and DPO, now a standard method for aligning LLMs, started as a rotation project in his group. In 2024 his lab showed for the first time that a diffusion model could match an autoregressive model on text quality at GPT-2 scale. These signals are from Ermon's interview with Sarah Guo on No Priors.
- The Inference Bottleneck. Autoregressive models generate one token at a time, and you can't produce the 10th token until the first 9 exist. Ermon says that workload "does not map well to GPUs." It is memory bound: the chip spends most of its time moving weights through the memory hierarchy and does very little arithmetic. He calls this a fundamental problem of the architecture, and the reason Inception bet against it.
- The Transformer Precedent. In 2017 the field switched from RNNs to transformers because RNNs processed tokens sequentially and trained slowly. Transformers processed many tokens in parallel and scaled better for training. Ermon argues diffusion is the same move applied to inference: it processes many tokens at once, so the inference workload looks like the training workload. His summary of the bitter lesson: the more parallel solution eventually wins.
- Inference Sets the Economics. Ermon says AI economics come down to intelligence per watt and intelligence per dollar, and both are decided at inference. Reasoning models scale test-time compute. RL post-training is bottlenecked on generating rollouts so the model can explore and get scored. A model that scales better at inference therefore also scales better during RL post-training.
- The 2024 Proof Point. Ermon's lab trained a standard transformer as a diffusion model on the same data as an autoregressive baseline at GPT-2 scale, under 1 billion parameters. It reached the same perplexity with the same parameter count and generated text about 10x faster, because it outputs many tokens at once. That result is why he started Inception: to see what happens at commercial scale.
- Words Have No Midpoint. Diffusion was built for continuous data, where you can interpolate between 2 pixel colors and still get something meaningful. Between 2 words there is nothing in between. Extending denoising to discrete text and code took years of R&D and what Ermon calls "a new science" before it worked.
- Small-Model Quality at Higher Speed. Inception's Mercury models benchmark on par with Anthropic's Haiku, Google's Flash, and OpenAI's mini and nano models while running significantly faster. The company is about 2 years old, has around 50 people, and serves these models in production today. Diffusion LLMs can't run on vLLM or SGLang, so Inception built its own serving engine.
- Software Speed on Commodity GPUs. A voice-agent customer (rendered "Open Call" in the transcript) runs a pipeline of speech recognition, a reasoning LLM handling tool calls, and text-to-speech. It had been serving its LLMs on Cerebras chips to hit its latency targets. It switched to Mercury on Nvidia GPUs and got the same speed, with more availability, lower cost, and higher quality. Ermon says software gains multiply with hardware gains, so the two approaches stack.
- A Drop-In API. Inception kept everything backwards compatible: text in, text out, through an OpenAI-compatible API. The models follow instructions and produce structured JSON output. For the voice customer, Mercury did better than the model it replaced, so the customer kept its existing harness and safety layer unchanged.
- Steering From the First Step. An autoregressive model has to finish generating before you can check whether output meets a constraint, such as solubility for a generated molecule. Diffusion generates coarse to fine, so you can score and steer the object from the very first step using an external reward function or constraint set. Ermon points to academic evidence that diffusion models are easier to control in ways autoregressive models can't match. Inception is working out what product to build around that capability.
- The 20 to 30% Floor. Using OpenRouter's taxonomy of workloads, Ermon estimates that 20% to 30% of tasks are ones where latency really matters. He treats that as a lower bound for what diffusion models can address in the next 2 years. Fast models also create habit: "once you get used to a fast model, it's hard to go back," and buyers already pay more for faster versions of frontier models.
- The Closed-Stack Moat. Asked how a small company competes when big labs can absorb any new architecture, Ermon names 3 things: trade secrets in training, a serving engine nobody else has, and evals and data from real customers. Inception also built its own SFT, RLHF, and RL tooling and chose not to open source it. He admits the cost: less community contribution, harder adoption, and difficult on-prem deployment.
- Data Efficiency as the Next Bet. Speed was the first wedge because it was easy to test and measure. Ermon expects other capabilities to emerge, and cites academic evidence that diffusion models are more data efficient. Denoising works as data augmentation, since each training example is seen through many noisy views. If that holds at scale, diffusion could win the domains where data, more than compute, is the bottleneck.
Watch the full video at https://www.youtube.com/watch?v=N1rjtDs8blY.