The More Affordable Tokens Become, the Higher the Bills Climb: Why US AI Budgets Have Split Into Three Tiers and Why Everyone Is Turning to Chinese Models

Deep News
Yesterday

Under the cost pressure of tokens, American enterprise AI budgets are gradually taking on a three-tier structure. The first tier consists of giants building their own alternatives, the second tier consists of mid-sized firms cutting budgets or downgrading, and the third tier consists of small and medium-sized businesses that are simply turning to Chinese open-weight models. Behind this great wave of stratification, what US frontier labs like OpenAI and Anthropic are losing is both ends of the market: their largest enterprise clients are building their own tools, while their smallest enterprise clients are flocking to Chinese AI models. Earlier this year, a term was still popular in Silicon Valley: TokenMaxxing, meaning whoever burned more tokens had higher productivity, and budgets existed precisely to be burned. Yet this great token leap lasted only half a year, and the situation has now completely reversed; nobody talks about TokenMaxxing anymore. When the bills came due, companies of different sizes adopted completely different responses. American enterprise AI procurement thus split into three clear tiers, opening up a definite market space for Chinese AI.

Current reality: tokens got cheaper, yet they became unaffordable

The most striking number comes from Uber. According to foreign media reports, Uber burned through its entire 2026 AI coding budget in less than four months. It burned so fast because employees actively used AI as Uber required. In March, across Uber's five-thousand-person engineering team, Claude Code penetration climbed from 32% all the way to 84%, and a single engineer spent 500 to 2,000 dollars per month on tokens. An even more extreme case is the agent framework OpenClaw behind the shrimp-farming craze; one estimate holds that a single instance running autonomously for a full day can consume API costs equivalent to one thousand to five thousand dollars, while users paid only a 200-dollar-per-month Claude Max subscription for it; by early April, more than 135,000 OpenClaw instances were running out in the wild. Anthropic therefore quickly cut off this route in early April, requiring usage of such third-party frameworks to be billed separately, and some users' expenses rose tenfold or even dozens of times. Sixteen days later, GitHub halted new registrations for Copilot Pro and Pro+, and Vice President of Product Joe Binder gave an equally blunt reason: agentic workflows often consume more compute than a user pays for in a month, and the cost of a handful of requests can exceed the subscription price. There is a trend here that is widely misread. On one hand, the unit price of tokens is indeed falling, dropping by roughly an order of magnitude every eighteen months, but on the other hand, token consumption per task is rising faster than the price is falling. The result, surprising to many, is this: the cheaper the model, the more expensive the bill. Gartner expects global AI spending to reach 2.5 trillion dollars this year, up 69% year on year; at the same time, however, it expects 25% of 2026 AI budgets to slip into 2027. "Slipping" is the respectable term for cutting budgets. Separately, Gartner also expects that only 28% of AI infrastructure projects have fully delivered on their original business case. In other words, nearly three-quarters of the money burned failed to prove it was worth burning.

First tier: giants that are self-sufficient

The first tier consists of AI giants capable of building their own models; for them, saving money and promoting their own products can be accomplished at the same time. After Microsoft replaced Claude Code, what it pushed engineers to use was its in-house MAI-Code-1-Flash. This is a small model of roughly five billion parameters, released at this year's Build conference and rolled out across all GitHub Copilot tiers; a later version, MAI-Code-1.1-Flash, claims a further 73% cost reduction. Behind these moves, Microsoft has reportedly cut its internal budget forecast for Anthropic usage by more than a third, whereas the previous estimate was at least one billion dollars per year. Microsoft's position here is actually quite delicate. In November 2025, together with Nvidia, it invested 15 billion dollars in Anthropic, pushing that company's valuation to 350 billion, and also put Claude on Azure for external sale. But investing in Anthropic and cutting Claude budgets internally are not contradictory. For Microsoft, Claude is merchandise, while MAI is its own land. Meta was the most frenzied AI giant in the earlier TokenMaxxing era, even treating token usage as a performance metric. Under performance pressure, its employees once consumed 60.2 trillion tokens through Anthropic's tools in thirty days. At the peak, Meta internally estimated it would spend up to 10 billion dollars a year on Anthropic models. But starting at the end of June, it restricted employees from using Claude Code and OpenAI's Codex, with the applied AI group the most strictly limited, and some tasks even requiring an approval process. Their substitutes are MetaCode and Muse Code. According to media reports, the number of Claude users inside Meta has already dropped from about sixty thousand to thirty thousand, a figure inferred externally from token consumption and spending estimates. Even after the cuts, Meta's monthly Claude spending is reportedly still in the hundreds of millions of dollars. There is another reason Meta tightened Claude usage, namely model distillation. If code written with Claude enters its own training data, then Meta's models may absorb the characteristics of a competitor's model along with it.

Second tier: mid-sized companies can only cut budgets

The second tier consists of mid-sized companies like Uber: several thousand engineers, bills on the same order of magnitude as the giants, but no frontier model of their own to switch to. Uber CTO Praveen Neppalli Naga admitted in April this year that after the company deployed Claude Code to about 5,000 engineers, it burned through its entire annual AI budget in just four months, catching him off guard. The intent of Uber's senior leadership was to promote and popularize AI tools among engineers, and from that angle they achieved their goal. Data from May this year showed that 95% of Uber engineers used AI tools every month. But leadership had not considered the other side: the API cost generated by a single engineer each month reached 500 to 2,000 dollars. Facing an astonishing token budget, even a giant like Uber with a market value of 150 billion dollars had to urgently formulate strict tiered management, limiting employee usage, and meticulously calculating the cost of every token just as people once economized on paper. This has become a broad industry consensus. Box CEO Aaron Levie said he attended a dinner with chief information officers from Fortune 500 companies and found that the topic business leaders discussed most was not macroeconomic issues but their companies' token costs. Levie listed five ways CIOs are trying to cut budgets: assigning tasks to different models by workload, issuing agents of differing capability by user type, setting different spending caps for each team, requiring teams to justify the necessity of AI by use case, and some simply letting it run unchecked. After hitting this wall, this tier really has only three choices: cut the number of seats, downgrade to cheaper models, or push the budget into the next fiscal year. That 25% of AI budgets slipping, as Gartner noted, is mainly contributed by this tier. Since there are inexpensive and capable Chinese open-source models on the market, why don't these mid-sized companies simply switch? Because for American companies at this level, what they consider more is data security and compliance; they would rather use less than take on policy risk, since once regulators catch a flaw, the fine could far exceed the IT cost savings. The US government's attitude toward Chinese open-source AI has never been clear; how to regulate it, and whether to regulate it at all, remain unknown. On that open letter on open weights in July, signatories included Nvidia, Microsoft, Meta, IBM, Dell and Palantir; Anthropic did not sign, and its CEO Dario Amodei opposed a blanket ban and advocated targeted controls. And while the rules remain undecided, the legal departments of mid-sized companies usually choose to stand pat. Occasionally a big customer emerges from this tier as well. Airbnb chose Qwen rather than ChatGPT for its customer service scenario, currently the most presentable sample, and it also shows that once price appeal crosses a certain threshold, even a listed company's procurement process cannot stand in the way.

Third tier: startups flock to Chinese AI

The third tier consists of small companies and startups. They bear neither compliance burdens nor in-house development capability, so the only remaining decision variable is cost-effectiveness. For this tier, an AI census can be seen in the data from the model routing platform OpenRouter. Chinese models' weekly share on that platform has peaked at 46.4%, while American models account for 35.7%; in the single-provider rankings, DeepSeek leads at 17.6%, about 5.13 trillion tokens per week, Qwen stands at 13.9%, and Anthropic at 14.8%. The price gap explains almost everything. OpenRouter's Justin Summerville said Chinese open-source models are 60% to 90% cheaper than the flagships of Anthropic and OpenAI; on specific models, as of June this year, DeepSeek V4 Flash charged 0.14 dollars per million input tokens, while GPT-5.5 charged 5 dollars, with the overall price-gap range between 4 times and 100 times. It needs explaining that this 46% does not equal Chinese models capturing nearly half of the American enterprise market. OpenRouter's user base naturally skews toward developers and small and mid-sized teams, and the traffic of large clients that sign enterprise contracts directly with Anthropic and OpenAI is not in this pool. But the existence of the three-tier structure has already become an indisputable trend. Chinese models reaching 46% on OpenRouter shows that what they have captured is not the American enterprise market but the American startup market. For US frontier labs, something worse lies in the future: they are losing their next generation of customers. Startups that begin with DeepSeek today will likely not turn back three years from now when they have grown up.

OpenAI cuts prices twice in a row to expand share

Facing simultaneous pressure across these three tiers of enterprise AI budgets, OpenAI's response is to cut prices. And it cut them twice in a row. In the July 30 round, GPT-5.6 Luna was cut by 80%, to 0.20 dollars per million input tokens and 1.20 dollars for output, while Terra was cut 20% to 2 dollars and 12 dollars; in the September 22 round, GPT-6 Sol dropped to 2 dollars and 10 dollars, whereas the previous generation GPT-5.6 Sol was 4 dollars and 20 dollars, GPT-6 Luna fell to 0.10 dollars and 0.50 dollars, and cached input reads were discounted to a tenth. OpenAI made clear these are permanent prices, not promotions. OpenAI's own explanations for the cuts all point to internal efficiency: hardware routing optimization, inference software improvements, context caching improvements; they also mentioned that Sol autonomously rewrote a production kernel in a supervised test, cutting serving overhead by 20% and raising token generation efficiency by more than 15%. Although they do not mention competitors, the direct cause of these two rounds of price cuts, and the market positioning, are very clear. Anthropic's Claude Opus 5.5 went live about ninety minutes before OpenAI's September release, priced at 4 dollars and 20 dollars, exactly half of what GPT-6 Sol charges. OpenAI CEO Sam Altman said on X that GPT-5.6 Sol is half the price of Claude Fable 5, and that he is "happy to deliver at a quarter of the price." Of course, the sustained upward push of Chinese models in the low-end market is a real threat. One detail is quite telling: the metric OpenAI now emphasizes is "intelligence per dollar," because in the face of a price gap on the order of 0.14 dollars for Chinese AI versus 5 dollars for American AI, American vendors have already given up direct price comparison and can only emphasize their own intelligence performance. Under the cost pressure of tokens, American enterprise AI budgets have taken on a three-tier form. The first tier builds its own alternatives, the second tier cuts budgets or downgrades, and the third tier simply turns to Chinese open weights. What US frontier labs like OpenAI and Anthropic are losing is both ends of the market: their largest enterprise clients are building their own tools, while their smallest enterprise clients are flocking to Chinese AI models.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10