AI model price war intensifies as Chinese rivals force major US developers to slash costs

Deep News
Aug 04

The artificial intelligence sector has entered a new phase of competition. Leading global AI firms are significantly reducing their product pricing while simultaneously enhancing the capabilities of their entry-level models, a direct response to the challenge posed by next-generation Chinese systems such as Moonshot AI's Kimi K3 and DeepSeek's V4 Flash.

The most recent company to adopt this approach is OpenAI. It has cut the price of its foundational frontier model, ChatGPT 5.6 Luna, by 80% on a per-million-token basis, while reducing the cost of its mid-tier 5.6 Terra model by 20%. This move comes just over a week after Google launched its cheaper Gemini 3.6 Flash and 3.5 Flash-Lite models. While Anthropic has not directly cut prices, it has replaced its lowest-cost Opus 4.8 model with the more capable Claude 5.0 at the same price point. The escalating global rivalry is driving down the cost of AI capabilities, but it is also placing significant pressure on the profit margins of these large corporations. Over the past several months, numerous major AI firms have announced cost-cutting measures and have begun restricting the scope of AI technology usage, a trend that even includes companies like Elon Musk's xAI, which had previously been a strong proponent of AI development. Despite a growing number of employees using AI, current reports indicate that the expected boost in actual productivity has not yet materialised.

The differing development paths of the US and China in AI, to some extent, reflect the long-standing industrial advantages of each nation. US companies tend to invest vast sums of capital to push the boundaries of the most advanced technology. In contrast, Chinese developers, leveraging their strong manufacturing and engineering capabilities, are creating AI models that are lower in cost, more streamlined in scale, yet nearly on par in terms of high-end capability. DeepSeek delivered a major shock to Western AI developers in 2025, and the Kimi K3 had a similar impact again in 2026. Individually, these events were significant enough to put pressure on companies like OpenAI, Google, and Anthropic. However, in the preceding months, many enterprises that heavily use AI had been complaining about rapidly rising token costs. So, when new models appeared on the market that offered comparable performance at substantially lower prices, the entire industry was strongly shaken. Today, large AI companies can no longer simply compete by increasing model parameter counts and the scale of training data. They are now forced to engage in genuine price competition, and OpenAI has already dramatically lowered the prices of several of its models.

It is noteworthy that OpenAI has not reduced the price of its most powerful flagship model. In fact, the price for the faster version of this model has actually increased. Nevertheless, its frontier-level models are now at historic lows in price, and the speed at which AI capabilities are improving while prices are falling is astonishing. In March of this year, OpenAI released ChatGPT 5.4, a model with more powerful agentic capabilities, priced at $2.50 per million input tokens and $15 per million output tokens. Now, GPT 5.6 Luna costs just $0.20 per million input tokens and $1.20 per million output tokens. This means that from launch to a significant price reduction, a frontier model has taken less than four months. Following this cut, Luna's price is now within the competitive range of DeepSeek V4, the professional version of which costs $0.435 per million input tokens and $0.87 per million output tokens. The more capable GPT 5.6 Terra, after a 20% price reduction, is now priced at $2 per million input tokens and $12 per million output tokens, undercutting the high-profile Kimi K3, which is priced at $3 per million input tokens and $15 per million output tokens. Meanwhile, GPT 5.6 Sol retains its price of $5 and $30 per million input and output tokens, respectively. OpenAI has even increased the price for its top-tier model: the 5.6 Sol Fast mode charges $10 per million input tokens and $60 per million output tokens, offering the same intelligence with lower latency to directly compete with flagship frontier models like Claude Fable 5 and Mythos 5.

The key question, however, is whether these price cuts can actually generate profit for OpenAI. Earlier this year, after OpenAI announced it was effectively abandoning plans to own its own first-party data centres, its lease agreements with "Neoclouds" became even more critical. For instance, a massive $300 billion computing resource procurement agreement with Oracle has become a vital foundation for OpenAI's continued operations. However, if there were any previous doubts about how OpenAI could afford such costs, that question is now even more pronounced. OpenAI is currently sustaining ongoing losses on its subscription business and missed key revenue targets earlier this year. In 2025, despite continued revenue growth, the company still lost tens of billions of dollars. OpenAI has committed to spending approximately $600 billion on computing resources by 2030. Even if revenue continues to grow, it may fall far short of covering such enormous costs. The further price reductions on its most popular, lowest-cost models mean that the company's profit margins could shrink dramatically or even vanish entirely. This may partly explain the external discussions about Nvidia potentially providing OpenAI with a $250 billion investment to support its operations.

Of course, OpenAI is not the only company facing this problem. Over the past year, Google's spending on AI infrastructure has been roughly nine times the revenue from its cloud business. Anthropic's recent ability to achieve an annualised revenue profit has primarily come from a limited agreement to cheaply lease the Colossus data centre from xAI. The cost of running and building AI has not suddenly decreased, yet companies are continuously lowering prices while providing faster and more powerful models at a lower cost. On the surface, these figures seem illogical. The AI industry often invokes the "Jevons Paradox" to explain the accelerating adoption of AI. In Jevons' time, more efficient coal engines did not reduce coal consumption; instead, they led to a surge in total coal demand as their usage expanded. Today, AI companies argue that as the cost of using AI falls, people will discover more applications, driving a significant overall increase in AI usage. This logic likely underpins the current strategy of token price reductions. If tokens are cheap enough, users will increase their usage substantially, leading to higher overall revenue through volume growth. Furthermore, by the end of this year, as a new generation of AI accelerator chips becomes more widespread in data centres, the token processing capacity per unit of power could increase tenfold, making even lower-margin businesses profitable through scale. Nvidia's future Vera Rubin platform is also noteworthy. Nvidia claims that this platform will deliver another 10x improvement in token performance efficiency. In theory, the advantages of the Blackwell and Vera Rubin platforms could make AI inference servers a more profitable industry. Even so, it is difficult to imagine that large AI companies can cover their current massive investments solely from AI business revenue, especially given the intensifying competition and the continued decline in token prices.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10