Builders Free to read

Memory Bandwidth Is the Next Scaling Law

Fractile bet on fast inference chips in 2022, before most people knew what inference meant. Walter Goodwin on why fast chatbots are faster horses.

· 4 min read

Walter Goodwin on No Priors

Walter Goodwin is the founder and CEO of Fractile, a full-stack AI chip company building very fast inference chips for the largest models. He started the company in the summer of 2022, betting early that AI would shift into a deployment era where test-time compute and speed decide what models can do. Fractile now has about 150 people spanning workload research, front-end design, physical design, backend implementation and advanced packaging. Its first DRAM-based platform is due to ramp in the second half of next year. These signals are from Walter Goodwin's interview with Sarah Guo on No Priors.

  1. The Inference Bet. Goodwin founded Fractile in summer 2022 on two observations: foundation models were generalizing across everything, and a few people were arguing that models needed far more compute at test time. He spent about two years after that explaining what inference meant. One inference chip company that recently went public was a training chip until about 18 months ago, he says. He points out that marginal cost is the cost you pay every single time you run a model.
  2. A Zoo of Identical Chips. Google's TPU, Meta's MTIA, Microsoft's Maia and OpenAI's in-house chip are structurally similar, Goodwin says. They are delivered through a small number of ASIC houses, with Broadcom the largest at around $2 trillion in market value. Nearly all use HBM memory, tensor cores for matrix math and TSMC's advanced packaging. Very few teams go all the way down the stack, from architecture to the physical design work that has you "flying back and forth to Taiwan or Korea every week."
  3. Owning the Whole Stack. Most chip projects hand front-end code to a partner like Broadcom, which turns it into a GDS2 layout file for TSMC and owns key analog and process-specific work. Fractile keeps all of it in house with about 150 people, "skinny" in every function. Goodwin says that lets the company run a closed loop instead of waiting on a partner. In AI chips, more than any earlier chip market, you have to place the right bet and then strike fast.
  4. From SRAM to DRAM. For its first two years Fractile built an SRAM-based chip, similar to Groq or Cerebras, because on-die SRAM gives enormous bandwidth and thousands of tokens per second. In late 2023 and 2024, growing context lengths made Goodwin nervous about how that design would scale. Fractile then ran skunkworks projects with memory vendors and its foundry partners to get very high bandwidth to cheaper, higher-capacity DRAM. The resulting platform combines GPU-like capacity with Groq- or Cerebras-like speed.
  5. Fast Chatbots Are Faster Horses. Goodwin borrows Henry Ford's line: a snappier chatbot is the "faster horses" version of fast inference. The bigger gain comes from running multi-trillion-parameter models at many thousands of tokens per second, which makes long-running agents radically faster. Today's fast inference chips have high-bandwidth memory with very low capacity. They can't run long-context attention and hand that work back to GPUs.
  6. Cost per Gigabyte. When you serve thousands of users at data center scale, Goodwin says inference economics collapse down to the cost per gigabyte of memory you deploy. Chips need bandwidth high enough to reload model weights and state thousands of times per second. They also need memory cheap enough to hold huge models and long contexts. Fractile's core technical goal is getting very high bandwidth out of the world's lower-cost DRAM.
  7. A Rolling Frontier of Bets. New models arrive about every two weeks, and the exact form of attention in frontier Chinese open-source models changes every couple of weeks. Goodwin anchors on what stays constant: every new LLM wants orders of magnitude more memory bandwidth. Fractile keeps a set of architectural bets ready and triggers the ramp on whichever fits. If you can carve out a lasting 3 to 6 month lead, he says, "you will be winning all of those deployments."
  8. Physics Still Sets the Clock. Goodwin wants AI to shrink front-end design from about 12 months to a fraction of that, but Amdahl's law applies: the steps you can't speed up become the limit. Foundry turnaround is 3 to 5 months even on a rush lot. A chip needs a 3 to 5 year amortization window and a 12 to 18 month volume ramp to make financial sense. He doesn't expect new chips every few weeks; faster cycles give more shots on goal, and Nvidia systems already combine six to nine custom chips that a single-chip startup must answer.
  9. Divide It by Four. A CEO of one of the top three semiconductor companies told Goodwin it would take 10 years for AI to turn an architect's intent into a usable GDS2 file. Goodwin's rule for the past couple of years is to take any timeline estimate and divide it by four, then trim a bit more. The hard part is place and route, where conventional algorithms for NP-hard problems run for days. Fractile is working on rough approximations to iterate faster, while leaving final sign-off to Cadence and Synopsys tools that confirm a design is clean under foundry rules.
  10. Think Before the Experiment. As AI handles more of the thinking in chip design and in AI research itself, the expensive physical or compute-bound experiment becomes the bottleneck. Goodwin cites Beren Millidge, who suggested on Dwarkesh Patel's podcast thinking for the human equivalent of 100 years before firing off an experiment. That makes reasoning tokens before each costly step close to a "moral compunction." It is also Fractile's workload bet: fast chips make that extra thinking cheap.
  11. Scaling Laws for Bandwidth. Fractile's chip targets about 25 times more bandwidth per chip than an HBM-based chip. Over the past 20 years, compute has grown roughly a millionfold while memory bandwidth grew about 40x. Mixture-of-experts models would save large amounts of compute by getting much sparser, say 1 in 128 or 1 in 256, but that is too bandwidth-hungry to serve on today's GPUs. Goodwin argues that with more bandwidth, labs can spend fewer flops for the same intelligence, and new chip properties can pull model design toward them.
  12. Why Labs Won't Bet the Farm. Goodwin jokes that the main purpose of first-party chips today is to lower what labs pay Nvidia, since those chips are architecturally similar. Labs face an asymmetric risk if they go all in on proprietary silicon. If a rival finds a 5x efficiency gain that only works on its own hardware, a lab "could die in the 9 months" it takes to deploy enough of that chip. That keeps room for third-party chip makers offering new capabilities, especially since open models like Kimi push frontier labs to compete on speed as well as weights.

Watch the full video at https://www.youtube.com/watch?v=OpeCP4wCxkA.

sig·nal·ful /ˈsɪɡ.nəl.fəl/ adjective — full of signal.

Get Signalful in your inbox.

One story free among every issue. Members unlock all, plus access the full archive.