Put a Hoop in Your House
Speechify bought its own GPUs because engineers renting compute were afraid to use it. Cliff Weitzman compares it to Michael Jordan's backyard hoop.
· 5 min read
Cliff Weitzman is the co-founder and CEO of Speechify, the text-to-speech platform used by more than 60 million people. He is dyslexic and learned to read when his father read Harry Potter aloud to him; he moved to the United States at 13 without speaking English. At Brown he built a tool that read his textbooks aloud, and he credits it with getting him through college. Speechify now accounts for 98% of text-to-speech app installs, has served more than 770 billion words, and is moving into business APIs and voice agents with its own models. These signals are from Cliff Weitzman's interview with Harry Stebbings on 20VC.
- A Hoop in the House. Speechify bought its first rack of Nvidia GPUs in 2022 because engineers on rented compute rationed it, worried about costing the company tens of thousands of dollars. Weitzman and his brother Tyler compared it to a young Michael Jordan paying $20 an hour for court time when he needs a hoop at home. Owning the hardware removed that hesitation, and the models got better. The company now spends tens of millions of dollars on GPUs, including multiple 72-card racks of Nvidia's Rubin chips.
- The 1.5x Math. An H100 costs about $30,000 to buy, while renting one from a hyperscaler runs $3.50 to $5 an hour, or $35,000 to $50,000 a year. Renting for a year costs about 1.5 times owning, and Weitzman expects the hardware to keep working well past its 3-year warranty, possibly 10 years. Large training runs also need memory sitting next to a big cluster, which is hard to rent without a multi-year commitment. He compares the return on spare cash in GPUs with a long bond yielding about 5%.
- Layered Capacity. Speechify sizes compute against a seasonal curve: if November is 100% of demand, October is 140% because of back to school and December is 80%. Weitzman buys about 20% of baseline outright, signs long-term hyperscaler contracts for another 25%, and rents spot instances for the rest. Older chips move to inference, where a cheaper GPU still returns audio in about 100 milliseconds; the company still runs some work on K80s. The newest chips go to training, where every minute of waiting is lost ground against competitors.
- A Floor Under GPUs. Weitzman says GPUs depreciate more slowly than people assume because they sit in clean, cooled, constantly maintained rooms. He cites a recent Nvidia deal with Blackstone, BlackRock, Apollo and Goldman Sachs to buy back GPUs for up to 25% of their value if a borrower defaults. That gives lenders a floor and should lower borrowing costs, the same move Elon Musk made by getting banks to finance Solar City panels. He calls last year's Oracle and OpenAI deals "ridiculous" but rejects the circular-economy fear about Nvidia, because a GPU's teraflops have real value anywhere in the world.
- The Hidden Work of Hardware. Nvidia doesn't sell to startups directly, so Speechify buys through vendors like Dell, which Weitzman calls a GPU rack supplier more than a PC company. When a French supplier ran weeks late, he threatened to switch, because a late delivery means paying data center rent on empty space. Speechify pays premiums of around $100,000 to jump delivery queues, insures trucks carrying "multiple houses worth of GPUs," and bought a liquid cooling sidecar because most data centers aren't set up for Rubin racks. Energy, he says, is now the biggest constraint.
- Compute as Headcount. The turning point came when Weitzman saw talented engineers moving at a fraction of their possible speed because they lacked compute to match their ideas. Speechify has 45 engineers and wants 150; its strongest engineers get a dedicated DGX rack each, and about 25% of the team was waiting for capacity. Each engineer runs 5 to 18 agents on long-horizon experiments, and the job comes down to about 10 good decisions a day. One of his best engineers builds synthetic data sets instead of models.
- The ElevenLabs Mistake. Weitzman met the ElevenLabs founders in London around 2022, admired them, and passed on their strategy because he expected text-to-speech APIs to become a commodity that runs on any phone. He calls it "the biggest strategic mistake I made in the history of Speechify." ElevenLabs used the API as a wedge, adding voices, cloning, speech-to-text and conversational agents sold to CEOs and CIOs, and now sells to governments. He took the lesson: offer an excellent product cheaply or free, then keep adding to it.
- Back in the Race. Stebbings argued that Speechify should stay a consumer company instead of fighting ElevenLabs and Sierra for third place. Weitzman says AI has made his engineering team about 10 times more productive, so the extra capacity should go to business customers as well. Anthropic spent years behind OpenAI and Facebook started behind Friendster, so second place can change. Speechify's new API, Simba 3.2, costs $10 per million characters against $100 at ElevenLabs and $196 for OpenAI, by his numbers. "The best way to lose is not to be in the race."
- Credit Only in Production. Speechify ignores token leaderboards and gives credit only when a feature reaches users without bugs. Weitzman compares it to milk delivery: a bottle left down the road spoils, so the job ends at the customer's door. On one call he turned his laptop camera toward his phone, used a voice prototype live, recorded the bugs and sent 14 notes. A 19-year-old engineer fixed all of them that night while waiting for 3 training runs to finish.
- Token Discipline. Weitzman encourages heavy token use but will let people go who burn tokens for no reason, such as 15,000 tokens on a tiny feature. A non-engineer friend built an app in 2 hours and then scraped 25,000 pages of a site when 20 would have done. Weitzman points to an Anthropic paper in which a model spent $12,500 of tokens over 2 weeks to train a better model, which he counts as money well spent. Following Claude Code creator Boris Cherny, he says the work is the loop: set the target, define how to measure it, iterate until you hit it.
- Hiring for Slope. Weitzman disagrees that seed startups face the hardest hiring market ever, since the labs' $15 million-a-year packages target a different pool. When Speechify had 21 people, 18 had been a CEO, CTO or VP of engineering. Today he hires math olympiad winners, Kaggle champions and physics graduates who may never have coded, because AI lets smart, hungry people get very good within 6 months. Interviews ask candidates to build something that passes unit tests and to change a large open-source codebase using agents.
- Data, Compute and Good Questions. Weitzman's brother has had a rare autoimmune neuroinflammatory disease for 6 years. Weitzman took a blood sample every week for 15 weeks, ran genome, protein and RNA analysis, and matched it against daily quality-of-life data on his GPU cluster in Scottsdale, Arizona. He is now buying a $5,000 pocket sequencer to collect genomes from other patients and plans to design candidate molecules with AlphaFold. He says GPUs already helped locate the lesion in his father's prostate cancer.
Watch the full video at https://www.youtube.com/watch?v=hj5oRzAnp2M.