OpenAI and Anthropic adopt "cost-per-task" metric for better AI cost evaluation

Deep News
Aug 14

The battle over AI pricing is intensifying, leading OpenAI and Anthropic to jointly reshape the industry's underlying pricing metric. They are shifting from a "cost per token" framework to a "cost-per-task" model, aiming to redefine the value proposition of their premium offerings for enterprise clients.

Over the past few weeks, both US AI leaders have publicly explained this shift. In a July blog post, OpenAI CFO Sarah Friar argued that while low-cost models have cheaper tokens, achieving good results "may require more attempts, more time, or more human review." Anthropic's Head of Financial Services, Jonathan Pelosi, echoed this in a recent interview, stating that token cost is only a "proxy metric" for measuring AI compute usage. He emphasized that "for every enterprise, the better proxy is the actual cost to complete a task."

This change in pricing narrative comes as competitors like Meta and Space offer models at extremely low token prices, squeezing the market. According to data from enterprise spending management platform Ramp, corporate adoption of OpenAI and Anthropic has slowed in recent months, with existing customers increasingly turning to cheaper open-source alternatives.

Third-party data delivers an even more direct signal: Anthropic's most powerful and expensive widely available model, Fable 5, accounts for only 11% of its Claude software spending. Based on this, Ramp concluded that it has "found a new ceiling for what enterprises are willing to pay for AI."

From "Token" to "Task": Resetting the Pricing Anchor

A token is a unit of data processed by an AI model and a common way to measure workload. The industry has long assumed that a lower cost per token makes software more economical for customers. OpenAI and Anthropic are now working to overturn this consensus.

Sarah Friar wrote: "Lower-cost models may have cheaper tokens, but getting good results might require more attempts, more time, or more human review; more capable models have more expensive tokens but can complete the same task in one go."

Pelosi's view aligns: tokens only measure how much AI is used, not the output. He believes the cost per token is "only useful as a proxy metric" and that enterprises should instead focus on the actual cost of completing a task.

Low-Cost Pressure: From "Tokenmaxxing" to Cost Ceilings

The metric shift occurs amid increasing competitive pressure. Multiple vendors are pushing for lower token costs, aiming to undercut companies like Anthropic that offer stronger models at higher prices. For example, DeepSeek's flagship open-source model, released in April, has input and output token costs that are a fraction of Anthropic's most advanced products.

Gil Luria, Head of Technology and Research at D.A. Davidson, offered an analogy: "US labs are trying to convey that using open-source tokens is like using cheap toilet paper — it does get the job done, but you need more, and there are other unpleasant side effects."

Companies that once encouraged employees to maximize AI usage through "tokenmaxxing" are now setting limits after experiencing "bill shock." Luria noted, "For many of these companies, internal costs have become prohibitively high."

Premium Model Value Questioned: Third-Party Data Reveals a Ceiling

Although vendors like Anthropic publish their own per-task cost estimates, a growing number of customers are turning to independent benchmarking sources like Artificial Analysis and Vals AI for objective assessments. Rayan Krishnan, co-founder and CEO of Vals, pointed out, "When labs report their own results, they often use internal benchmarks, making true comparisons impossible."

On a cost-per-task basis, models from Anthropic and OpenAI remain the most expensive on the market, with Anthropic's Claude Fable 5 topping the list. However, both companies also offer more competitively priced options, such as OpenAI's GPT-5.6 Luna.

The strongest blow to the high-premium narrative comes from Ramp's data: Anthropic's most powerful and expensive widely available model, Fable 5, accounts for only 11% of Claude software spending. "With Fable 5, we have found a new ceiling for what enterprises are willing to pay for AI," Ramp wrote in its report. "Beyond this point, higher performance is not worth the price."

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10