> ## Content Index
> Fetch the complete content index at: https://www.signalful.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# The Watt Is the Denominator
- URL: https://www.signalful.com/the-watt-is-the-denominator/
- Published: 2026-10-06T15:49:00.000Z
- Updated: 2026-10-07T05:15:40.000Z
- Description: Google will spend more than $200 billion on infrastructure this year. Amin Vahdat runs it, and he says chip specs are the wrong measure.
- Author: Signalful Editorial
- Tags: Builders, Google, Amin Vahdat, AI Infrastructure, Energy, #sf-0050, #yt-bGph8GwB3Sk, #Import 2026-10-06 15:50

Amin Vahdat runs AI infrastructure at Google, a job he was named to at the end of 2025, which puts him in charge of TPUs, data centers and the networks that connect them. The hosts call it the biggest capex buildout in history, with Google alone expected to spend more than $200 billion on capex this year, most of it on data centers. Before AI infrastructure, Vahdat spent years leading Google's data center networking work, and earlier he was a computer science professor at UC San Diego. He works side by side with Google DeepMind and says he talks with Demis Hassabis or Koray Kavukcuoglu several times a week. These signals are from Amin Vahdat's interview on Sequoia Capital's Training Data podcast.

1. **Delivered Goodput.** Vahdat calls FLOPs and other chip-centric numbers theoretical maximums, and says Google holds itself to delivered goodput: the useful work a real workload completes under real failure conditions. Performance rarely comes down to one chip; it depends on how 2, 16 or 10,000 chips work together with the CPUs that feed them and the network that joins them. His example is solving a problem on paper: if a mistake sends you back to step 1, you're still doing work (throughput), while goodput is the total time it took to reach the answer. Goodput started as a Google term, he says more of the industry is picking it up, and Google measures it per watt.
2. **Failure as a Constant.** At 100,000 accelerators, something fails multiple times a day and, depending on configuration, multiple times an hour. Training, serving and agent workloads are synchronous, so one dead chip can halt the whole job while Google finds it, restores a checkpoint and restarts. There is no single common cause to fix; Vahdat calls it "a long tail of constant discovery" across networking, hardware, compilers, runtimes and the models themselves. One defense is optical circuit switching, where MEMS mirrors redirect light between fibers, so Google can swap a failed TPU rack for a spare in milliseconds without moving any fiber.
3. **Doubling Every Six Months.** Google has to double its effective serving capacity, measured as the hardware's ability to generate tokens, roughly every six months. Vahdat says as much or more of that gain comes from software as from hardware, through "dozens or hundreds" of model and runtime improvements landing one after another. In Google's experience, the model side delivers most of the gains in intelligence per watt. Hardware still adds 2x or more year over year, a multiplier he says everyone higher in the stack can count on.
4. **Purpose-Built Buildings.** A traditional data center is a 25- to 30-year building that has to host hardware lasting about 6 years, so it stays general enough for many generations. Google now co-designs AI data centers with the hardware that goes in them, down to cooling and power distribution. A storage rack draws 10 to 40 kilowatts while a TPU or GPU rack draws hundreds, so a row built for 30 storage racks looks nothing like a row built for 1 or 2 AI racks. Vahdat says designing for full fungibility over 30 years means building "too big and too overbuilt."
5. **Agents Need CPUs.** Long-horizon agents, which took off at the start of 2026, remove the human who used to set the pace of requests to the model. The gap between calls drops from seconds or tens of seconds to milliseconds, and between calls a CPU parses the response and pulls context from local DRAM, another machine's DRAM, SSDs or hard drives. Vahdat says demand for CPUs, networking and storage is going "through the roof" alongside demand for accelerators. Google now has to choose between mixing CPU racks into dense TPU buildings or putting them next door and paying for it in network cost, reliability and hundreds of microseconds of latency.
6. **Two Chips for 2026.** About two years ago, Google debated whether its 2026 generation should be one chip or two, and it shipped two: TPU 8i for inference and TPU 8t for training. The call turned on market share. If inference were 2% or 5% of the market, a chip 2x faster at it wouldn't justify the fixed cost of a new program, but at a projected 30% to 60% over the chip's lifetime, it did. Vahdat says each chip can still run the other's workload, so Google doesn't have to predict the exact mix across a six-year hardware life.
7. **Primitives Since TPU v1.** In 2013, Vahdat says, the smartest people held that you don't build a custom accelerator for one workload when Moore's law keeps doubling general-purpose performance, and some people inside Google doubted the bet. The first TPU served translation and voice recognition, the second added training, and transformers and recommender systems for ads followed. Vahdat says the architecture, at a medium level of detail, hasn't really changed since v1: large matrix multiply units, a SparseCore for vector and scatter-gather operations, and remote loads and stores over the ICI network. That small set of primitives has carried many generations of models, which is why Google can plan chips two to five years out.
8. **Delaying a Tape-Out.** Google always has five or six chip generations in flight, from production through bring-up, tape-out, design and concept, and DeepMind works with the hardware team on all of them. When DeepMind finds a model change that would run much faster with hardware support, engineers from both teams spend days or weeks in the same room, and the result is often a model tweak that gets 90% of the benefit. For a big enough win, Google will delay a tape-out by a week or two. Vahdat says researchers know a 0.5% gain isn't worth stopping a tape-out, and that this kind of change in flight would be "somewhere between hard and impossible" across company boundaries.
9. **AI Designing Hardware.** Google uses Gemini to design hardware for future Geminis. Vahdat says his hardware engineers now use as much AI, by token count, as his software engineers, and both the time from design kickoff to tape-out and the time for bring-up are shrinking. Data center planning, such as choosing between a gigawatt building and a 100 or 200 megawatt campus, used to run on spreadsheets and people. AI now pulls the information together for the humans who decide, though he says it isn't replacing their judgment yet.
10. **Power as the Binding Constraint.** Vahdat says every input is a hard constraint, but power is the most fundamental, since the others look solvable with time. Google prefers to connect to the grid, gives utilities years of notice, and pays for new transmission lines and substations so other customers' rates don't go up. If Google needs a gigawatt in 2028 and the utility can supply 700 megawatts, it can fill the gap with its own solar and batteries, then send power back to the grid during peaks like the hottest two weeks of the year. Going it alone at 99.99% reliability would mean building about 2 gigawatts to get 1, which is why he says sharing with the grid means "everybody wins."
11. **Seven-Year-Old TPUs.** Vahdat says Google's 7- and 8-year-old TPUs are still at 100% utilization, a remark that got more pickup than he expected. TPUs depreciate over about 6 years, and Google replaces whole pods, which can hold around 9,600 chips, once newer generations' power efficiency makes the swap worth it. Old training clusters can't cover inference demand, because training gets concentrated on one continent while serving has to run close to users everywhere. Serving clusters are smaller and mix storage, compute and accelerators, and he says they aren't necessarily cheaper per megawatt.
12. **Orbit and the 2036 Rack.** Google is pursuing orbital data centers as a moonshot: in a sun-synchronous orbit, solar panels get about 1.4x the power and sunlight 98% to 100% of the time, against roughly 30% on land, which mostly takes batteries out of the equation. Cooling and repairs get harder in space, and links would run over free-space lasers instead of fiber, but Vahdat sees "no fundamental showstoppers." On the ground, he expects the 2036 frontier rack to be built centrally and packed far more tightly, with a small fiber bundle and possibly multiple megawatts in a single rack. Crews would wheel it in, connect water, power and fiber, and turn it on.

Watch the full video at [https://www.youtube.com/watch?v=bGph8GwB3Sk](https://www.youtube.com/watch?v=bGph8GwB3Sk&ref=signalful.com).