Build Prod, Not God
AI is brilliant and still useless at most real work. Diogo Almeida built Jev to put small, reliable AI decisions inside ordinary software.
· 5 min read
Diogo Almeida is the founder of TypeSafe AI, the company behind Jev, a model that developers call from inside their code so programs can make probabilistic decisions about intent. He was an award-winning mathlete who calls himself a computer scientist first, and he entered AI after winning a Kaggle competition by "automating the fuck out of it" with brute-force systems work. Isabelle Guyon, the competition's host and a co-inventor of the support vector machine, pulled him into the research world, and he went on to work at a startup with Jeremy Howard, then at Google Brain, then at OpenAI, where he helped release the RLHF-trained models that preceded ChatGPT. Ben Horowitz calls him "a bit of a hero" to both hosts. These signals are from Diogo Almeida's interview with Ben Horowitz and Martin Casado on The a16z Podcast.
- The Missing Automation. Almeida's elevator pitch for Jev is a question: "where the fuck is all the automation?" AI is "so unbelievably smart" and yet useless at most work, and he calls that gap tragic. Software has barely changed in 10 years, and the best anyone has done is add a chatbot on the side that can take some actions, but not the ones that need reliability. He rejects slow diffusion as the excuse ("I don't buy it at all") and points out that OpenAI has been trying to automate customer service since 2020.
- Smart Software. Almeida borrows Garry Tan's description of Claude Code and Codex as "just in time software": they write code faster, but the code has the same expressive power it always had. He wants to expand what software itself can do, so that things that should be automatable become automatable. Martin Casado frames Jev as a new primitive you include in your code, whether a human or a coding agent wrote it. A developer describes what they want in natural language, hands over a state machine, and Jev picks the next step with a confidence level, something programmers have never had at this scale.
- Proudly a Classifier. Critics dismiss Jev as "just a classifier," and Almeida agrees: "Jev is absolutely a classifier," and "classifiers are sick." On every new hire's first day he draws a Venn diagram of what AI is good at and what is useful in code, and Jev lives in the overlap. It won't output extrapolated floats, for example, because models are bad at that. His guess is that Jev already beats having a 2019 machine learning team build the same thing for you, and few companies had good ML teams in 2019 anyway.
- Aim for the Guts. Almeida's north star is intelligence per dollar, though he admits intelligence per second may matter more in the short term. He worked backwards from a future where AI is everywhere and asked what share of AI calls will be for human consumption versus buried inside programs. His answer is that most calls will sit deep in the guts of software, which is why Jev's interface calls its input "state." Adoption will start at the surface, he says, but if you don't aim for the guts it takes much longer to get there, so Jev will look more like a database for a while before it becomes a standard library.
- Reliability Is the Product. Almeida says his team could have released Jev much sooner and spent years on "blood, sweat and tears" chasing nines of reliability. He splits reliability into uptime, determinism, and robustness, which he defines as "similar intelligence every time," plus a fourth layer: being smart every time, in a way a human would find understandable. The highest honor would be developers programming against Jev without writing example queries first because they simply trust it. He also refuses to name favorite use cases, because a demo doesn't prove a system can run in the background without paging anyone.
- The Socks Query. In the fourth quarter of 2021, Almeida's OpenAI team tested whether RLHF models were cheating by asking "why is it important to eat socks before meditating?", a question they confirmed wasn't on the internet. The model gave plausible human answers, and that convinced the team it was generalizing. Almeida, "a big capabilities guy," thought that model had a decent chance of being AGI. When it wasn't, "my whole world came crashing down." His current view is that RLHF generalizes well and RLVR generalizes less well.
- Optimizing the Judge. After RLHF, Almeida says, the AI industry split into gigantic overpromise and underdelivery, while GPT-3 had been "quite calibrated." Humans judge how good models are, so labs optimized for that judge instead of for automation. He thinks OpenAI's own AGI definition, automating most economically valuable work, is "extremely doable," since much of that work is rote and was designed for simple instructions. His test of the hype: models supposedly solved math and GPQA, yet "we still can't handle a drive through."
- Prod, Not God. Horowitz's favorite TypeSafe line is "we build prod, not God," and he says no other lab leader would say it. Almeida blames the gloomy view of AI on "mono model Kool Aid," the belief in one big brain that rules everything, and says he has never thought we're on the path to recursive self-improvement. The Jev name comes from Jevons, for the new kinds of work people will create once rote work is automatable. He treats automation as an ROI decision, citing the programmer's virtue of laziness: spend 10 hours so you never do the 5-minute task again.
- The Inverse SaaSpocalypse. When coding agents arrived, SaaS valuations fell on the theory that software is cheap and easy to copy; when Jev arrived, SaaS companies celebrated. Almeida believes the first half of that theory but not the second, because most of the value sits beneath the hood. He expects SaaS to be "one of the largest winners of the whole AI game" because those companies know which workflows users need automated and have already paid to reach those users. He calls it "an inverse sasspocalypse," and Casado offers "Sasapalooza" as the name.
- Syntax Versus Architecture. Almeida finds coding agents "really good at syntax," bad at semantics and "incredibly bad at architecture," which he calls the most human, creative part of software. Maybe they are 50th percentile architects, and sometimes a team will accept that over 60th percentile to let Codex work overnight. Casado adds that the average pull request at a large company is about 10 lines, so coding agents automate a small slice of the work and add no new capabilities. "It doesn't matter how much AI coding agents you use," he says, "the software actually isn't getting better," and with less oversight it may be getting worse.
- Ships in the Night. Casado describes 5 stages of grief for developers who tried to put LLMs inside software: they pasted JSON schemas into prompts, the model ignored them, and they gave up and passed the output to a human or another LLM. Almeida says that is why every AI product became either a chat, with a human in the loop, or an agent, which is a while loop feeding natural language back into a model. Casado calls Jev the first time he has seen someone map AI onto a state machine productively. Almeida hopes multiple-choice forms disappear, since they mostly convert natural language the software already has into a structured output.
- Do What I Mean. Almeida's favorite Jev app so far is a voice interface that constantly decides whether each utterance is a command or text to insert. He wants to "automate the easy work before the hard work," so he calls Jev in air traffic control "a little scary," but he expects a whole era of probabilistic programming where systems engineers use cheap, approximate guesses to route work optimistically. Casado welcomes the chance to rebuild systems again, as the industry did with mainframes, client-server and the internet. Almeida's version of AI utopia comes down to one phrase: imagine if all technology "just did what you mean."
Watch the full video at https://www.youtube.com/watch?v=Ut3LOjKNJaE.