The Inference Foundry: How Baseten Turned the GPU Shortage Into an $11B Company

Every AI product ships on top of the same shortage: not enough GPUs, in not enough places, priced by an auction nobody fully controls. Baseten built the plumbing that hides that shortage from the rest of the industry. In eighteen months, four funding rounds turned that plumbing into a company priced between eleven and thirteen billion dollars.
The bottleneck nobody put in the pitch deck
We built machines that predict the next word in a sentence better than most humans can. We built machines that can look at a photograph and flag what a doctor should worry about, and machines that watch and understand video the way 12 Labs taught AI to watch. None of that has anything to do with silicon, except that it has everything to do with silicon.
Every one of those models runs on a GPU, and every GPU is a physical object: fabricated in Taiwan, shipped, racked, cooled, and rationed across data centers that cannot be built as fast as demand for them grows — a constraint extreme enough that StarCloud is trying to skip it entirely by putting the data center in orbit. Nvidia’s Blackwell chip packs 208 billion transistors, 2.5 times the density of the generation before it, part of how Nvidia turned scarce silicon into the arms dealer’s advantage of the AI boom. None of that transistor count matters if the chip is sitting in the wrong data center on the wrong continent when a customer’s traffic spikes at 2 a.m.
That is the problem Baseten sells the fix for. Not a better model. A way to find, provision and route the physical compute a model needs, in under five minutes, across more than twenty clouds, so the shortage never reaches the person typing into a chat window.
Four engineers who broke first
Baseten was founded in San Francisco in 2019 by four people who had already lived the problem from the inside: Tuhin Srivastava (CEO), Amir Haghighat (CTO), Phil Howes (Chief Scientist) and Pankaj Gupta. Srivastava, Haghighat and Howes had all worked together at Gumroad, the online marketplace for creators, where Haghighat ran engineering and Srivastava and Howes were data scientists pulled into building the company’s fraud detection and content moderation systems. None of them had signed up to be full-stack infrastructure engineers. The job made them into one anyway, and it showed them exactly how much of a data scientist’s time gets eaten by deployment rather than modelling. Gupta came from a software engineering role at Uber.
They launched Baseten commercially in April 2022, disclosing $20 million already raised: an $8 million seed round co-led by Greylock and South Park Commons Fund, plus a $12 million Series A led by Greylock. The seed round’s backers included OpenAI co-founder Greg Brockman, Figma co-founder Dylan Field, and DeepMind co-founder Mustafa Suleyman, a strange detail to find attached to an infrastructure company almost nobody outside machine learning had heard of yet, and a useful signal of how early some people had already priced in the shortage.
Baseten, by the numbers
| Founded | San Francisco, 2019, by Tuhin Srivastava (CEO), Amir Haghighat (CTO), Phil Howes (Chief Scientist), Pankaj Gupta |
|---|---|
| Origin | Three co-founders from Gumroad, the fourth from Uber |
| Apr 2022 | $20M disclosed (seed + Series A), led by Greylock |
| Feb 2025 | $75M Series C at an $825M valuation, days after DeepSeek’s R1 release |
| Sep 2025 | $150M Series D at a $2.15B valuation, led by BOND |
| Jan 2026 | $300M Series E at a $5B valuation, anchored by a $150M Nvidia check |
| Jun 2026 | $1.5B Series F, split-priced between $11B and $13B |
| Core system | Multi-cloud capacity management: thousands of GPUs provisioned in under five minutes, across 20+ clouds |
The moment the industry noticed the bottleneck
For most of Baseten’s first five years, inference was the boring half of AI, the part that happened after the interesting work of training a model was done. That changed in January 2025, when the Chinese lab DeepSeek released R1, a reasoning model that matched frontier performance at a fraction of the compute cost anyone had assumed was required. The message the industry took from it was blunt: efficient inference was not a rounding error on the AI budget. It was the budget.
Baseten’s Series C landed three and a half weeks later: $75 million led by IVP and Spark Capital, valuing the company at $825 million. CNBC’s own headline on the round called it what it was, a raise coming “following DeepSeek’s emergence.” The company did not cause the moment. It was positioned exactly where the moment landed.
What Baseten actually sells
Strip away the marketing language and Baseten is a logistics company for a scarce physical resource. Its core system, multi-cloud capacity management, treats GPU capacity the way a freight broker treats container ships: as inventory scattered across more than twenty clouds and regions, none of it owned outright, all of it needing to be found, priced and routed to wherever a customer’s traffic is spiking, in minutes rather than the weeks a hyperscaler contract usually takes to provision — the same physical-power constraint that has Crusoe Energy feeding AI compute without waiting on the grid, and is driving a new generation of micro-nuclear reactors built to power the AI buildout.
In its own case study with Nvidia, Baseten reports serving DeepSeek-V3 and DeepSeek-R1 with 38% faster latency and 225% better cost-performance than the prior generation, and a 60% throughput increase for Writer’s Palmyra models running on the same infrastructure. None of those numbers come from a smarter model. They come from squeezing more usable work out of the same finite chips: custom kernels, decoding strategies tuned per model family, and a routing layer deciding which of its more than twenty clouds gets the request.
The customer list reads like a cross-section of companies that discovered, the hard way, that a working model and a production system are not the same thing: Writer, Abridge, OpenEvidence, Clay, Zed, Gamma, Sourcegraph and Bland all run inference through Baseten rather than build the orchestration layer themselves. That is the same lesson every hardware-adjacent infrastructure company eventually teaches its customers. Building the thing that works is not the same job as running the thing that works, at scale, at 2 a.m., across a shortage you do not control.
$11–13B
Series F valuation, June 2026
Four rounds in eighteen months. Each one priced faster and higher than the last, on the back of measurable workload rather than press coverage.
Four rounds in eighteen months
Trace the rounds and the shape of the last eighteen months becomes visible. September 2025: $150 million Series D, led by BOND, at a $2.15 billion valuation, bringing total funding past $285 million. Jay Simons of BOND joined the board. Four months later, January 2026: $300 million Series E at a $5 billion valuation, more than double the Series D price, with Nvidia itself writing a $150 million check alongside lead investors IVP and CapitalG. Five months after that, June 2026: a $1.5 billion Series F, split-priced between $11 billion and $13 billion depending on the investor, led by Altimeter Capital, Conviction and Spark Capital, co-led by Sands Capital and Wellington Management.
Baseten’s own announcement called it the company’s “fourth fundraise in 18 months,” and put a number on why investors kept coming back faster each time: revenue grew 20x and inference volume grew 40x over the year leading up to the Series F. Total funding, across every round since 2019, now runs past two billion dollars.
20x / 40x
Revenue growth / inference volume growth, trailing 12 months
The number that got Nvidia to write its own check. Growth measured in workload, not narrative.
The valuation trajectory is aggressive by any standard, and this piece does not endorse an eleven-to-thirteen-billion-dollar price on an infrastructure layer sitting on top of GPUs it does not own. What is verifiable is the pattern behind the price: growth measured in inference volume, not press coverage, and Nvidia putting its own money into the company selling the software that gets the most out of Nvidia’s own chips.
Founder lessons from the inference layer
Necessity produces founders faster than ambition does
Srivastava, Haghighat and Howes did not set out to build infrastructure. They were a marketplace’s engineering team, drafted into fraud detection and content moderation because the company needed it, and the drafting is what taught them how much of a data scientist’s week gets swallowed by deployment rather than modelling. The company they built afterward is a direct translation of a problem they lived, not a problem they researched. If you want to find the real gap in a market, look for the team that got stuck solving it for somebody else first.
Sell the moment the market discovers your pain
Baseten spent five years selling a problem most of the industry considered secondary. DeepSeek’s R1 release in January 2025 forced a re-pricing of that assumption industry-wide, and Baseten’s Series C landed three and a half weeks later. The company did not manufacture that moment and could not have timed it. What it could do, and did, was already be the answer when the question changed. Build the capability before the market agrees it matters, so that when the market changes its mind, you are not starting from zero.
The scarce resource is the business, not the model
Baseten does not train frontier models and does not compete with the labs that do. It competes on a narrower and, it turns out, more defensible question: who gets the GPU, from where, in the next five minutes. Owning the orchestration layer over a resource nobody can manufacture fast enough is a different business than owning a model that a better model can obsolete next quarter. When everyone is fighting over the same scarce input, the company that allocates the input can end up more durable than the companies competing to use it.
Capital compounds when growth compounds
Four rounds in eighteen months, each one arriving faster and priced higher than the last, is not a fundraising strategy. It is what happens when 20x revenue growth and 40x volume growth show up in the data room before the round is even open, and when a strategic investor like Nvidia is willing to put $150 million of its own balance sheet behind the company reselling access to its chips. Growth that is measurable in workload, not narrative, is the only kind of growth that keeps pulling capital back faster than the founders ask for it.
Final word
The interesting AI story is rarely the model. It is the shortage underneath the model: the chips that cannot be fabricated fast enough, the data centers that cannot be built fast enough, and the five minutes it takes, or does not take, to find a GPU when a customer’s traffic spikes. Baseten’s bet is that whoever solves the allocation problem outlasts whoever wins any single model generation. Eighteen months and four funding rounds is not proof that bet is right. It is proof that, for now, the people supplying the chips agree it might be.
Sources
- Baseten Secures $150M Series D as the Premier Inference Platform for AI’s App Layer, BusinessWire, September 5, 2025
- Baseten (company blog), Announcing Baseten’s $150M Series D, September 2025
- Pulse2, Baseten: $150 Million Series D Raised At $2.15 Billion Valuation For AI Inference Platform, September 2025
- AI startup Baseten raises $75 million following DeepSeek’s emergence, CNBC, February 19, 2025
- Cooley, Baseten Announces $75 Million Series C, February 2025
- Baseten (company blog), Announcing Baseten’s $75M Series C, February 2025
- Bloomberg, AI Inference Startup Baseten Raises $300 Million at $5 Billion Valuation, January 20, 2026
- SiliconANGLE, AI inference startup Baseten hits $5B valuation in $300M round backed by Nvidia, January 20, 2026
- BusinessWire, Baseten Raises $300M at a $5B Valuation to Power a Multi-Model Future, January 23, 2026
- Baseten (company blog), Announcing Baseten’s $300M Series E, January 2026
- BusinessWire, Baseten Raises $1.5 Billion to Power the Next Era of AI Inference, June 22, 2026
- Baseten (company blog), Announcing our Series F, June 2026
- TechCrunch, AI inference startup Baseten reportedly raising $1.5B months after its last mega-round, June 18, 2026
- Baseten nabs $20M to make it easier to build machine learning-based applications, TechCrunch, April 26, 2022
- GlobeNewswire, Baseten Gives Data Science and Machine Learning Teams the Superpowers They Need to Build Production-Grade Machine Learning-Powered Apps, April 26, 2022
- Nvidia, Baseten Scales AI Inference With NVIDIA Blackwell GPUs (customer case study), nvidia.com
- Baseten (company site), Meet the engineers behind Baseten, baseten.co/about-us