AI Inference
The economics and infrastructure of serving models at scale — chips, runtimes, and margin structure.
Companies in this report

Fireworks AI — "the most bang for GPU," built by the team that ran 50 trillion daily inference instances at Meta — rides open-weight inference's surge from 1% to 30% of tokens. But buyers see few differences across providers and switch easily, threatening the edge the whole bull case rests on.

Baseten — "best-in-class for model deployment," with customers like Cursor and Notion — reports 400% net dollar retention and zero churn among its top 30, riding open-weight inference's rise. But buyers multisource and switch easily, raising the question of whether that retention holds as inference commoditizes.

Together AI scaled from $44M to $1B ARR in under two years by owning its GPU fleet rather than renting — a vertical-integration bet that aims to capture margin and availability beyond what asset-light peers typically offer as open-weight adoption accelerates. The open question is how that owned compute fares if AI demand cools.

Modal hit an estimated $300M ARR in roughly eight months by collapsing messy, heterogeneous AI workloads into a single Python decorator — positioning it more as a hyperscaler competitor than a Baseten one. The open question is whether that pull holds as GPU supply normalizes.

RadixArk is the commercial company behind SGLang, one of the three leading open-source LLM inference engines, serving trillions of tokens daily across xAI, Google, Microsoft, and dozens of other production deployments. With a $100M seed round backed by three competing chip vendors and top-tier AI angels, RadixArk is betting it can convert massive open-source adoption into a durable commercial business — a thesis with both powerful tailwinds and structural headwinds.