Harvey, a legal tech unicorn, saw its gross margin drop to negative 50% within months after lawyers began using AI agents at scale, and only returned to profitability after rolling out an in-house model built on China’s Kimi, according to an opinion article by lawyer Lin Shanglun published by BlockTempo.
Lin writes that the case shows how token consumption from professional users has already gone far beyond what Silicon Valley initially expected, putting the traditional flat-fee subscription model under pressure in the agentic era.
Harvey, described in the article as a global legal tech leader with a valuation above $10 billion, has long focused on top-tier law firms. It charged enterprises and law firms steep seat fees of roughly $1,200 to more than $1,500 per person per month, usually under an all-you-can-use annual contract structure common in software.
From the outside, that kind of pricing would normally suggest software-style margins above 80%. But Lin cites financial details disclosed by Bloomberg and other foreign media showing that Harvey’s gross margin slid from a highly profitable level to negative 50% in just a few months, leaving the company in a position where serving each additional customer meant taking a heavy loss.
The article says the shock did not come from rising API prices. Major foundation model providers have been cutting API rates over the past two years. The problem, Lin argues, was that professional users increased their token consumption far faster than anyone had expected.
AI agents changed legal workloads
In the early stage of legal AI adoption, lawyers mainly used these systems for single-turn Q&A, occasional translation, or short memorandum summaries covering only a few pages. The computing load for each task was limited.
That changed once multi-step, autonomous agent workflows entered the field. Lawyers started feeding entire batches of litigation files running to hundreds of pages, as well as cross-border M&A contracts that could reach thousands of pages, directly into the system. To complete deep due diligence or rigorous contract cross-review, the AI had to run large amounts of chained reasoning in the background, repeatedly call external tools for retrieval, compare internal inconsistencies, and rewrite its own output multiple times.
According to Lin, that pushed token consumption for a single task up by dozens or even hundreds of times.
He also argues that this kind of reliance is hard to reverse. Once lawyers experience reviewing more than a thousand pages of contracts in a matter of minutes, they are unlikely to return to fully manual review. As heavy users with strong budgets kept sending large volumes of work to agents around the clock, Harvey’s fixed monthly fee structure was quickly overtaken by ballooning downstream API costs.
Harvey’s cost response
Facing what the article describes as a survival crisis driven by inference costs, Harvey moved to launch its own Harvey Tenet model architecture. Lin says that shift cut inference expenses by several times and helped move the company’s gross margin from negative territory back into profit within a short period.
He points to one detail as especially striking: Harvey, a U.S. legal tech company whose funding included leadership from the OpenAI Startup Fund, turned to China’s Kimi as the foundation for its cost-saving model rather than relying on a U.S. language model.
What Lin says the case reveals
Lin draws two industry observations from the episode. First, compute demand in top professional use cases has moved well beyond the early assumptions made by Silicon Valley investors and analysts, with usage growth arriving fast and proving difficult to reverse. Second, AI competition between the United States and China has shifted away from surface-level political rhetoric and into a direct contest over cost and performance.
In Lin’s telling, if a top U.S. unicorn needs a Chinese model to win the cost battle and restore profitability, the competitive fight has already moved to the underlying economics of model deployment.
Questions raised in the article
Why did Harvey’s margin fall to negative 50% despite high pricing?
Lin says Harvey charged fixed seat fees under all-you-can-use annual contracts, but once lawyers adopted agent-based workflows, token consumption per task surged by dozens of times, and API costs ate through subscription revenue.
What industry signal does the margin squeeze send?
The article argues that demand for computing power in professional scenarios is much higher than Silicon Valley had expected. It also says Harvey’s move to Kimi for cost control shows that U.S.-China AI competition has entered a phase defined by cost and efficiency.

