Harvey, the legal AI startup valued at $15.6 billion, spent the first half of 2026 losing money on every additional query its customers ran. Gross margins that sat near 50% in January had cratered to roughly minus 50% by June, Bloomberg reported on Monday, after a March update to Harvey’s agents sent customer usage soaring and token consumption on OpenAI and Anthropic’s usage-based pricing jumped twentyfold. The fix was not a price increase. It was a new model, built on Chinese open weights.

Margins turned positive again in August, when Harvey shipped Harvey Tenet, an in-house model post-trained on Moonshot AI’s open-weight Kimi K3 with Fireworks AI. Bloomberg, citing a person familiar with the company, describes the recovery as a direct consequence of moving Harvey’s flagship workloads off frontier-lab APIs. Neither Harvey nor Bloomberg has disclosed the post-Tenet margin figure, or what share of the hardest legal work still routes to GPT and Claude.

The arithmetic behind the collapse is simple. Harvey sells annual per-seat licences with unlimited usage inside each seat, while paying its model providers per token. A partner who runs one query a week and an associate who leaves a diligence agent grinding through a data room overnight pay the same. Over the period that token usage rose twentyfold, annual recurring revenue roughly doubled, from about $190 million at the start of the year to more than $400 million. Revenue doubling while costs rose twenty times is a ratio no flat price survives. At a 50% gross margin every revenue dollar carries 50 cents of cost to serve; at minus 50% it carries $1.50.

Harvey’s own engineering posts make the motive explicit. A 20 August post on Tenet states that “open-weight models have cheaper per token prices,” and reports the post-trained model delivering better answers on Review Tables at roughly one-tenth the cost per cell and cutting cost per query on Firm Knowledge by 90%. Fireworks’ companion write-up gives the mechanism on Harvey’s Legal Agent Benchmark: Tenet costs $5.92 per task against $5.62 for base Kimi K3, but lifts the all-pass rate from 10.8% to 19.7%, pushing the cost per fully completed task from about $52 down to about $30. In an earlier June study, Harvey and Fireworks ran open-weight GLM 5.1 as the primary agent with Claude Opus 4.7 called as an occasional advisor; the hybrid passed 18 of 100 tasks for $368 while Opus alone passed 14 for $954. As the Fireworks post put it, “The frontier model shows up as a callable tool, not as the dependency the product is built on top of.”

Harvey is the loudest example but not the only one. Bloomberg names Abridge, which is building a clinical documentation model on Nvidia open weights, Decagon, which now handles about 80% of customer support inquiries with its own models, and Ramp, which has been weighing an in-house model since raising $750 million in June. Rogo and Canva are reportedly on similar paths. Investors including Sequoia Capital and General Catalyst are backing the trend, on the logic that owning the model attacks one of a startup’s largest cost lines while reducing dependence on OpenAI and Anthropic.

Ramp co-CEO Karim Atiyeh told Bloomberg the calculus has flipped quickly. “It made absolutely no sense a year ago. It’s starting to make a lot more sense now,” he said. Not everyone is convinced. Matt Kraning, a partner at Anthropic investor Menlo Ventures, cautioned that building a model requires specialised staff and higher upfront costs, and that for many companies the exercise is closer to marketing than engineering. “In most cases, it tends to be a lot of cosplay,” he told Bloomberg.

Why it matters

For two years the working assumption in venture circles was that application-layer AI companies were captive customers of the frontier labs, and that OpenAI and Anthropic could price enterprise tokens knowing their downstream apps had no exit. Harvey’s minus-50% quarter is the clearest public evidence yet that the exit exists, and that it can be executed inside a single quarter by a company with roughly 3,000 customer organisations and a fresh $550 million round. Every vertical AI startup selling seats against a metered model bill now faces the same three levers: meter usage, cap it, or swap the model. Cursor and GitHub Copilot chose to meter and passed the cost to customers. Harvey chose to swap, and by its own measurements got a better product for it.

The geopolitics are unavoidable. Harvey counts the OpenAI Startup Fund among its investors and previously built on closed US models from OpenAI, Anthropic and Google; its flagship now sits on weights released by a Beijing lab. That will sharpen the Silicon Valley debate over restricting Chinese open weights, and it complicates the revenue story at the labs: Anthropic has reportedly shown investors a slide noting that Harvey still needs Opus for its hardest tasks, which is true but also a concession that the frontier model has become a specialist tool rather than the platform.

The caveats matter too. Every quality number here comes from Harvey or its training partners, measured on Harvey’s own benchmark, and the best configuration still completes fewer than one task in five. Harvey’s customers now run on a router whose thresholds are invisible to them, which is why enterprise buyers are starting to ask for model-change notice clauses in their renewals.

What to watch

Harvey’s next ARR or margin disclosure will show whether the August recovery holds as usage keeps climbing, and whether the company follows through on the new pricing model it was reported in July to be developing. Watch, too, for any enterprise token price cut from OpenAI or Anthropic in response; either lab could make the open-weight arbitrage far less compelling overnight. Abridge, Decagon and Ramp have yet to publish margin numbers of their own, and Moonshot’s next Kimi release will test whether a Chinese lab can keep the capability gap narrow enough for US startups to keep betting their flagships on it.

“It made absolutely no sense a year ago. It’s starting to make a lot more sense now.”
— Karim Atiyeh, Co-CEO, Ramp
-50%
Harvey gross margin by June, down from ~50% in January
20x
Increase in Harvey token usage in 2026
$15.6B
Harvey valuation after $550M round this month
90%
Reduction in cost per query on Firm Knowledge after Tenet, per Harvey