đ Iâm Ivan. I study how top 1% startups grow.
Startup Riders is brought to you by PitchMagic.ai (built by yours truly!):
Investors rarely tell you why they passed. PitchMagic does, before you hit send.
I built it because friends kept asking me for feedback on their pitch decks. PitchMagic grades yours on the 4 questions every investor asks (do I get it, is it real, can it be huge, is it them?), scores it out of 100, marks the exact words to change on every slide, and ranks the 3 fixes that matter most to get your foot in the door.
Pro members get 10 full reviews free (âŹ49 value) â get your code.
Hello there!
This week Iâm diving into Fireworks AI, a company running open AI models for other companies (i.e. Cursor, Uber, Notion) that recently raised $1.5B at a $17.5B valuation (4 years after its founding, another AI-wave rocket-ship):
In a nutshell:
Product: think about it like an âengine roomâ for AI. Companies that donât want to depend on OpenAI / Anthropic take a free open model (i.e DeepSeek or Llama) and Fireworks runs it for them (fast, cheap and inside their product). Today it apparently already handles 40T+ tokens a day (tokens are the chunks of words AI reads + writes) for 10K+ companies.
Money: the model is whatâs quickly becoming a classic = pay per use (i.e. $0.30 per million tokens in, $1.20 per million out on a model like DeepSeek), renting chips by the hour (i.e. $8/hr for an Nvidia H100), or pay to train your own custom model. Theyâve already passed $1B in annualized revenue as of July this year with around ~200 people ($5M per employee!).
Driver: loved how they frame it here: âCompanies are no longer renting general intelligence. Theyâre building their own.â
8 growth levers in this drop:
Sold Metaâs internal speed tricks to everyone
Made switching painless at the moment costs typically start to hurt
Had every hot new free model running on launch day
Built custom for the fastest-growing customer + then sold it to everyone
Made engineers part of the âunofficialâ sales team (inside the customerâs Slack)
Made a custom model cost the same as a generic one
Turned big clouds into both suppliers and a sales channel
Bet on deciding which model runs for each task
đ Quick note on editorial + methodology: this analysis focuses on the 80/20 mechanics that explain their growth (itâs not a comprehensive profile, not an endorsement or investment advice). I use AI like a fund leverages an analyst for groundwork, the direction + judgement are mine. Company-reported figures are marked as such, treat directional estimates as directional.
Zero to one
Who: 7 engineers who used to run the AI plumbing at my good old employer Meta + Google (6 from Meta and four of them actually from the PyTorch team aka the free toolkit most of the world uses to build AI models today).
Lin Qiao (CEO): led the PyTorch there growing the team from 5 to ~300 people, before that IBM and LinkedIn. She started Fireworks at age 48 (after shelving a 2015 startup plan to learn how to lead people at Facebook). She meant to stay âone year or twoâ and stayed 7.
Dmytro Dzhulgakov (CTO): from Ukraine whoâs a PyTorch core maintainer after 11 years at Meta, before intern at Google.
Benny Chen: a new Zealand-born engineer who ran ads infra also at Meta.
Plus Dmytro Ivchenko, James Reed, Pawel Garbacki (all ex-Meta) and Chenyu Zhao (ex-Google).
Where they started (fall 2022): they all apparently left Meta about 2 months before ChatGPT came out and the first plan was a platform to help companies use PyTorch. Benny, in 2024 said: âExactly how, we didnât really figure out completely... ChatGPT wasnât out yet, so we had to pivot somewhere in the middle.â They raised $25M from Benchmark and Sequoia before having let alone launching a product. Apparently, fun detail, they started in a building literally called the âSequoia buildingâ before Sequoia invested.
The wedge: PyTorch was of course free and open so companies like Walmart, Disney, Tesla, Netflix etc used it too and they kept coming back to Linâs team asking âcan you build this training platform for us? Can you build this serving platform for us?â Lin took it as a signal the industry was ready / market timing was right, âand then more importantly, we have the key.â
The MVP: a platform that launched in August 2023 (11 months in). Run and customize free models through a simple web connection with a free tier for devs:
First users: âa lot of ex-coworkers from Meta,â says Benny plus teams who trusted the PyTorch name (this is a classic Silicon Valley dense network kindling effect), and by mid-2024 the list already included Cursor, DoorDash and Quora.
Growth Mechanics
Lever 1: Sold Metaâs internal speed tricks to everyone
âOur biggest differentiation is... Fireworks off the shelf is faster than both of the offerings and second is weâre building a system, not just a library.â (Lin, Sequoia)
What happened: at Meta theyâd built these tricks that let AI run at planet scale with things like custom code that squeezes more work out of chips, compression that makes models smaller without losing (much?) quality, etc. At Fireworks they then packaged those tricks + rented them to companies through a simple connection.
The details:
In Jan 2024 they said their engine ran 4x faster than open-source alternatives
Speed matters here, youâve probably experienced it using any of the models consumer facing interfaces: âlow latency is a critical part of product viability... people are not patient enough to wait for half a minute.â
By Series B in July 2024 they were serving already 140 billion tokens / day.
Meta itself became one of their customers this year.
But this edge doesnât last on its own as we see on Artificial Analysis (independent speed rankings) Fireworks sits 6th of 11 on DeepSeekâs newest big model.
So what: turning a big companyâs internal tools into a product is a classic founding play and in this case it got them in the door / great wedge. Speed was an advantage but it isnât necessarily a lasting moat (rivals caught up within 2 years, its also a little silly to talk about moats in this market considering how fast everything moves, but worth keeping an eye out). Interesting what theyâve done to retain customers (levers 4 to 6 later).
Lever 2: Made switching easy when costs start to hurt
âThey cannot open up the floodgate because theyâre going to scale into bankruptcy.â (Lin, theCUBE)
What happened: companies often build the first version of their AI products on OpenAI, Anthropic or similar because itâs the fastest way to test an idea. If the product takes-off usage can explode and so does the bill and that can be a problem. So Fireworks built for that moment or rather with that in mind. So they made it easy to move over with a few lines of code and the open models it runs cost much less to use.
The details:
Going bankrupt on AI bills comes up in 17 of the 44 interviews we went through for this deep-dive, itâs essentially their sales pitch.
It now reaches big companies too, the founder Lin recently said in a Sequoia interview: âtheir CFO is blocking their AI feature launch because of the costâ
The pain in the market has become so big / obvious that apparently they barely marketed. Lin recently said on 20VC âwe feel product speaks for itself... we didnât spend much time marketing at all.â
They apparently priced for the comparison table, moving in March 2024 to a flat price per token âbecause people would often make comparisons using output token costâ.
And just to show you how fast these switches are going for out there apparently Innovative Solutions (AWS partner) moved â90% of Anthropic inference spendâ to Fireworks in 2 weeks (case study, May 2026).
So what: they never fought OpenAI for the prototype. They waited for the moment the bill hurts and made leaving nearly free. In strategy terms theyâre the cheaper substitute for the closed labs, and a substitute only wins if switching is painless.
Lever 3: Had every new âhotâ free model running immediately on launch day
âWe are very proud of day zero launch always. Like we kind of become famous on day zero launch.â (Lin, Startup Grind Q&A, May 2026)
What happened: every time a new free model came out like DeepSeek, Kimi, Llama etc. etc., they tried to have it running for customers the same day.
The details:
DeepSeek R1 came out in Jan last year and they had a full breakdown 4 days later.
OpenRouterâs CEO says Fireworks âcaptured most of the inferenceâ on DeepSeek in the early days at least âas far as our metrics told usâ (Sequoia)
Two months after R1 they matched DeepSeekâs API price at $0.55 in and $2.19 out per million tokens.
They also know when to break their own rules intelligently for example with DeepSeekâs V4 release since it had bugs, so they held it back about 3 days.
They now host 200+ models and add one pretty much every week these days.
So what: the open-model boom to certain extent was luck in terms of a tailwind, but what they did differently was treat each launch as a free acquisition moment and try to never miss one (turning othersâ announcements into their own marketing).
Lever 4: Built custom for the fastest-growing customer, then sold it to everyone
âWe have finite engineers like everybody else. We would like prefer to have engineers make training more efficient... rather than like spin up like a inference effort.â (Federico Cassano, Cursor, Sequoia 2026)
What happened: Cursor needed its AI edits to feel instant 2 years ago, and Fireworks built it for them, then converted that âcustomâ work into a product line that they then started selling to everyone else.
The details:
They started working with Cursor when it was making about $2M a year. Lin said recently: âNow itâs a thousand times over the past three yearsâ
Cursorâs âFast Applyâ (the feature that writes AI edits into your code) ran at around 1K tokens a second which is about 13x faster than running a similar open model the standard way back in June 2024.
Lin says that work became a separate product line (FireOptimizer) once they saw other customers wanted the same thing.
Cursorâs own model (Composer 2) trained with Fireworksâ infrastructure and runs at around 6 to 10x lower cost compared to other frontier coding models.
On the other hand and also worth noting is the revenue concentration this represented, at one point it was around half of Fireworksâ revenue. The CTO says that was true âfor us and our competitors last yearâ though and no longer is the case (David Ondrej, Sept 2026). Also SpaceX acquired Cursor this August.
So what: building for the pioneer showed them what to package for the rest, offered a good opportunity to take a peek to look for what is coming around the corner. The price they had to pay was buyer power (one customer can squeeze you). The trade-off worked out for them so far.
Lever 5: Made engineers part of the âunofficialâ sales team (inside the customerâs Slack)
âYou think the Cursor founders want to go out to a steak dinner with me? Like they would laugh me out of the room... They donât even want to do Zoom calls. Like I had to do most of my communication with them in Slack.â (Bardia Shahali, VP Sales for Fireworks from Series A to C, The Crew Pod, Jun 2026)
What happened: their buyers were some of the best AI Eng in the world and a âclassicâ sales team would have probably struggled to persuade / connect. So engineers apparently did a lot of the selling. Trials ran as live problem-solving in a Slack channel shared with their potential customers (+ a salesperson in the background). Also for companies with no AI people of their own they send their researchers to do the first training runs with them.
The details:
Bardia the ex VP sales said: âone AE at Fireworks can maybe support 10, 20 million of ARR per year, but youâll need like 3 or 4 forward deployed engineersâ (in their case an AE closes the deal and the FDEâs work inside the project).
A trial runs â1 month, or maybe even 2 months,â with multiple Eng in the slack channel, âAEs are still around... but theyâre not driving that.â
The founders were the closers, Bardia says: âwhen they jump on a call with another engineer on the other side, the empathy just comes across... they have so much credibility.â
So what: when your buyers are engineers the trial is often the sales pitch so more often than not the people running it should probably be engineers too. It moves no big strategic force and anyone can copy the model of course but there is a bottleneck (a sort of cornered resource?) which is engineer talent and credibility. Meaning having engineers their customers' engineers look up to (i.e. starting with founders who built PyTorch).
Lever 6: Made a custom model cost the same as a generic one
âDo you want to feed the beast or do you want to control your own destiny? That was a very easy sell for them.â (Benny Chen, Boardroom Club Apex, Jun 2026)
What happened: a âcustom modelâ is basically a free model trained a bit more on your own data so it does your specific job better, normally a custom model needs its own chips so the cost goes up. They found a way to stack many custom versions on one shared base model so the extra cost dropped:
âown your AI instead of renting it from OpenAI.â
The details:
In September 2024 they launched serving hundreds of custom models âat the same cost as a single base modelâ.
Then Cresta (customer-service AI company) ran thousands of custom models at 100x lower cost than its previous GPT-4 setup.
The founders pitch it with simple math, where for example a company paying $100 a month for a closed model can train its own on Fireworks and run it for $10-20, plus a few dollars for the training, so it ends up paying 4-8x less.
Juicebox (recruiting search tool) for example went from spending $4M a year on AI to $800K after switching to models tuned for its use case.
Lin says 95% of their traffic now runs on customized models.
So what: customization can help build up switching costs. The customer owns the model (it's open) but the retraining loop aka the tuning on Fireworks' engine + engineers who know your setup don't come with. It's a potentially strong defence against the closed labs (though Baseten, Together and the big clouds all sell fine-tuning too, so call it a moat in progress?).
Lever 7: Turned the big clouds into both suppliers and a sales channel
âThe revenue for Fireworks is definitely a capacity constraint... itâs not a demand constrained environment.â (Benny Chen, The Information Bottleneck, Jun 2026)
What happened: they own no data centers. They rent chips from dozens of cloud companies, because their sales are capped by how many chips they can get (not by demand, for now). They also sell through the biggest clouds. Big companies often promise Microsoft a certain amount of spending (years ahead), and buying Fireworks through Microsoft's store counts toward that (no new budget to approve).
The details:
A multi-year deal with AMD (October 2025), with Nvidia and AMD both investors.
The CTO calls it a âvirtual cloudâ: âWe run on... several dozens of all this new clouds and traditional hyperscalers... from customer perspective you donât need to worry about compute procurementâ (David Ondrej, Sept 2026). To make that work they bought Hathora this March, a startup whose software moved video game servers between clouds automatically (now it does the same for AI work).
An AWS partnership (November 2025) and a launch on Microsoftâs AI marketplace (March 2026).
So what: when growth is capped by supply like their case the job is buying from as many sources as possible so no single supplier has leverage (i.e. Nvidia against AMD, cloud against cloud). Letting big companies pay with budget theyâd already set aside also removes the slowest part of an enterprise sale cycle aka gives them velocity.
Lever 8: Bet on deciding which model runs for each task
âThe frontier isnât a model. Itâs a router.â (Fireworks blog, Sept 2026)
What happened: in July they launched Nexus which sits in front of AI coding tools (i.e. Claude Code, Codex etc) and sends routine work to cheaper open models (keeping expensive ones for the harder tasks).
The details:
Early tests with Notion and Doximity apparently showed the product cutting about 30% the cost per accepted code change
It puts them head to head with OpenRouter on this which routes traffic across providers and was valued at $1.3B in May this year.
Fireworks runs its own company on it including from engineering to sales, finance and legal.
The brain inside Nexus is FireRouter, their first router model which came out literally 2 days ago where for each step of a coding task it picks between Claude Opus and two cheaper open models. In a month of internal testing it cost 57% less per session at 98.1% of Opusâs accuracy.
So what: if you pick the model you own the customer (to a certain extent), and the model makers become suppliers you can swap, early days on this bet though.
Thatâs it for this week friends!
Cheers,
Ivan
đ Building something in this space? We invest âŹ100K-3M at pre-seed and seed. If youâre raising or know someone who is, please send us your deck via DM.
đ Bibliography
Founder and operator interviews (44 interviews reviewed; key ones below)
Lin Qiao â Sequoia Training Data (Aug 2024)
Lin Qiao â Autopilot with Will Summerlin (May 2024)
Lin Qiao & Dmytro Ivchenko â Cognitive Revolution (Apr 2024)
Lin Qiao â Latent Space (Nov 2024)
Lin Qiao â Hanselminutes (Nov 2024)
Lin Qiao â Unsupervised Learning (Dec 2024)
Lin Qiao â 5 Year Frontier (Feb 2025)
Lin Qiao â The MAD Podcast (Mar 2025)
Lin Qiao â Turing Post (Aug 2025)
Lin Qiao â theCUBE + NYSE Wired (Nov 2025)
Lin Qiao â Super Data Science #971 (Mar 2026)
Lin Qiao â Startup Grind Q&A (May 2026)
Lin Qiao & Jensen Huang â GTC 2026 (May 2026)
Lin Qiao â Gradient Dissent (Aug 2026)
Lin Qiao â 20VC with Harry Stebbings (Jul 2026)
Lin Qiao â Candid with Index (Jul 2026)
Lin Qiao â RAISE Summit fireside and theCUBE at RAISE (Jul 2026)
Lin Qiao â Sequoia, Post-Training Is How You Keep Your Taste (Aug 2026)
Lin Qiao â NEW ECONOMIES (Sept 2026)
George Hu (President) â Adventure with Grace (Sept 2026)
Dmytro Dzhulgakov â David Ondrej podcast (Sept 2026)
Bardia Shahali (ex-Fireworks sales) â The Crew Podcast (Jun 2026)
Benny Chen â theCUBE at MongoDB.local (May 2024)
Dmytro Dzhulgakov â Sequoia AI Ascent panel (May 2025)
Dmytro Dzhulgakov & Federico Cassano (Cursor) â Sequoia Training Data (May 2026)
Benny Chen â Infinite Curiosity (Feb 2026)
Benny Chen â MindMakers (Apr 2026)
Benny Chen â Boardroom Club Apex (Jun 2026)
Benny Chen â The Information Bottleneck (Jun 2026)
Primary sources
Fireworks: Series D announcement (Jul 15, 2026)
Fireworks: Series C (Oct 28, 2025)
Fireworks: Series B (Jul 11, 2024)
Fireworks: platform launch (Aug 17, 2023)
Fireworks: Spring 2024 pricing update (Mar 1, 2024)
Fireworks: DeepSeek price match on Developer Cloud (Mar 18, 2025)
Fireworks: Innovative Solutions case study (May 5, 2026)
Fireworks: FireRouter with Opus (Sep 28, 2026)
Fireworks: FireAttention (Jan 8, 2024)
Fireworks: Cursor Fast Apply (Jun 23, 2024)
Fireworks: Multi-LoRA (Sep 18, 2024)
Fireworks: DeepSeek R1 deep dive (Jan 24, 2025)
Fireworks: AMD partnership (Oct 20, 2025)
Fireworks: AWS alliance (Nov 24, 2025)
Fireworks: Hathora acquisition (Mar 8, 2026)
Fireworks: Cursor Composer 2 (Jun 2026)
Fireworks: Nexus (Jul 2026)
Fireworks: pricing and serverless pricing docs (accessed Sept 26, 2026)
Index Ventures: Inference is the New Runtime (Oct 28, 2025)
Index Ventures: Candid with Lin Qiao (Jul 16, 2026)
Lightspeed: Our Investment in Fireworks AI (Oct 28, 2025)
Coverage and analysis
Contrary Research: Fireworks AI business breakdown (Aug 6, 2026)
Sacra: Fireworks AI (2026)
Latent Space: Why Compound AI + Open Source will beat Closed AI (Nov 25, 2024)
Turing Post: interview with Lin Qiao (Aug 23, 2025)
SemiAnalysis: Inference Race To The Bottom (Dec 18, 2023)
SemiAnalysis: Are Open Models Catching Up? (Aug 21, 2026)
Growth Unhinged: An AI-native guide for scaling from $1M to $10M ARR (Apr 12, 2026)
Eastwind: A Deep Dive on AI Inference Startups (Jul 10, 2024)
TechCrunch: Fireworks AI open source API (Mar 26, 2024)
TNW: Fireworks $1.5B Series D (Jul 16, 2026)
CNBC: Fireworks hits $17.5B valuation (Jul 16, 2026)
TechCrunch: SpaceX to acquire Cursor (Jun 16, 2026)
Baseten: Series F (Jun 22, 2026)
TechCrunch: Together AI raises $800M (Jul 1, 2026)
TechCrunch: OpenRouter valued at $1.3B (May 26, 2026)
Artificial Analysis: DeepSeek V4 Pro providers (accessed Sept 26, 2026)



















