a16z partner Sarah Wang said in a recent post that AI application-layer products should not copy model-layer token pricing. Instead, she argued, they should be priced around customer-recognizable units of work. Her core point is that token pricing ties application value to a unit whose cost keeps falling, while also making it hard for customers to predict context length, retrieval volume, or inference time. It can also push buyers to compare an application with raw compute in ways that do not reflect the product’s actual value.
Wang proposed a tiered pricing framework based on value. Under that structure, model layers can still charge by token, while application layers should charge by identifiable work units such as an account research brief, a code edit, or a completed query, potentially packaged through credits. In scenarios where attribution is clear and value is easy to define, she said products can charge directly by outcome, such as a resolved customer service conversation or a qualified lead. She cited Clay as an example, noting that its new pricing separates Data Credits from Actions and passes through inference-model cost fluctuations without markup.
According to ChainCatcher, a16z partner Sarah Wang said in a recent post that AI application-layer products should not follow model-layer token pricing. Instead, they should price around “recognizable units of work.”
In the post, Wang argued that token pricing anchors application value to a unit whose cost keeps declining. She also said customers often cannot reliably predict context length, retrieval volume, or inference time, and that token-based pricing can lead to inappropriate comparisons between an application product and raw compute.
Pricing by value tier
The article proposed a tiered pricing structure based on value:
- model layers priced by token;
- application layers priced by customer-recognizable work units, such as an account research brief, a code change, or a completed query, with credits used as a packaging layer;
- scenarios with clear attribution and clear value priced directly by outcome, such as a resolved customer support conversation or a qualified lead.
Wang added that credits should map to work of different levels of difficulty. That structure, she wrote, can help protect gross margins while separating “work value” from “delivery cost.”
Clay cited as an example
The post used Clay as an example. Under its new pricing, Data Credits, which cover third-party data, are separated from Actions, which cover orchestration work. For inference models with more volatile costs, Clay only passes through the cost and does not add a markup.
Wang said pricing tied to value rather than compute cost can help customers understand what they are paying for, while also helping product providers preserve margin.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.