tradingkey.logo
tradingkey.logo
Search

Why Wall Street Is Bullish on AI Chips After Kimi K3 Release

TradingKeyJul 24, 2026 7:00 PM

AI Podcast

facebooktwitterlinkedin
View all comments0

Moonshot AI’s Kimi K3, featuring a 2.8 trillion-parameter MoE architecture, has gained global attention for its high performance and competitive pricing. While the model's efficiency initially triggered concerns regarding reduced hardware demand, the Jevons paradox suggests that lower inference costs may spur exponential growth in AI adoption, ultimately sustaining or increasing demand for computing infrastructure. However, risks remain: if efficiency gains outpace application growth or commercial monetization fails, hardware investments could face downward pressure. Furthermore, intensified model competition and the shift toward custom silicon may challenge the pricing power of general-purpose GPU manufacturers.

AI-generated summary

TradingKey - Chinese AI large language models have once again become the focus of global market attention.

Recently, Chinese AI developer Moonshot AI officially released its latest-generation large model, Kimi K3. With its open-weights architecture, performance close to that of the world's leading models, and more competitive inference costs, Kimi K3 quickly sparked heated discussions in overseas tech circles and capital markets.

Following the announcement, the market briefly worried whether the logic of global tech giants continuing to invest hundreds of billions of dollars to build AI infrastructure would be impacted as more high-performance, low-cost AI models emerge, and the US AI chip sector also experienced a noticeable correction.

However, an increasing number of analysts believe that Kimi K3 may not necessarily be a negative for the AI chip industry. Instead, improved model performance and lower usage costs could lower the barrier to AI applications, bringing more users, higher call frequencies, and larger-scale inference demand. The "Jevons paradox" in economics may be playing out once again in the AI industry.

Why Kimi K3 Is Drawing Global Market Attention?

Kimi K3 is the flagship model launched by Moonshot AI in July 2026, primarily targeting application scenarios such as long-cycle programming, complex knowledge work, and AI agents.

According to disclosures by Moonshot AI, Kimi K3 adopts a Mixture-of-Experts (MoE) architecture, featuring 2.8 trillion total parameters and a 1-million-token context window, and natively supports visual understanding. Among its 896 expert modules, only 16 are activated per calculation to improve model capacity and inference efficiency. It is worth noting that 2.8 trillion is the total parameter scale, which does not mean that every generated token invokes all parameters.

Kimi K3 has performed outstandingly in frontend programming evaluations, once ranking near the top of the Frontend Code Arena leaderboard with a score of 1,679. However, this achievement is confined to specific programming tests and does not equate to comprehensively outperforming the flagship models of OpenAI and Anthropic in terms of overall capabilities.

Pricing is also a key reason for the attention surrounding Kimi K3. Official API pricing shows a rate of $3 per million input tokens—dropping to $0.30 upon a cache hit—and $15 per million output tokens. For AI agents that need to repeatedly retrieve codebases, enterprise documents, and long contexts, the caching mechanism can significantly reduce input costs.

However, a 'lower unit price per token' does not equate to a proportional decline in the actual cost of all tasks. Model output length, inference time, cache hit rate, and task completion rate will all impact the final cost.

Therefore, the real change brought by Kimi K3 is not simply launching a price war, but rather driving the broader adoption of advanced models with a higher price-to-performance ratio.

Why Low-Cost Kimi K3 May Increase Chip Demand?

After the release of Kimi K3, some investors worried that improved model efficiency would reduce GPU usage. However, judging from actual operations, a decline in unit inference costs does not directly equate to a drop in total chip demand.

Kimi K3 adopts a sparse Mixture of Experts (MoE) architecture, activating only a subset of expert modules during each inference, which helps control the computational overhead per run. However, the model still needs to store and schedule massive weight data, while processing contexts of up to 1 million tokens. In large-scale deployments, such tasks still rely on GPUs, AI accelerators, high-bandwidth memory, high-speed networks, and server clusters.

More importantly, as model costs decrease, the number of users and frequency of use may grow even faster. Following the launch of Kimi K3, user request volume rapidly approached the capacity limit of the existing computing clusters. Moonshot AI temporarily suspended some new user subscriptions to prioritize computing power resources for existing users. This indicates that as models become cheaper and easier to use, new demand can quickly consume the computing power saved by efficiency gains.

This is precisely the core of the 'Jevons Paradox': as the efficiency of resource utilization improves, the unit cost falls, but because the scope of application expands, its total consumption may actually increase instead.

In the AI industry, improvements in model efficiency can lower the barrier for enterprises to deploy intelligent customer service, coding assistants, office agents, financial analysis, and medical assistance systems. As more applications run continuously, the daily volume of inference requests will surge, ultimately driving continued growth in demand for computing power, memory, and data center infrastructure.

Low-Cost Models Do Not Mean Computing Power Investment Is Without Risk

While the Jevons paradox provides an optimistic explanation for AI chip demand, it is not an inevitable outcome. If the rate of model efficiency gains outpaces user and application growth in the long run, the decline in computing power required per task could still compress a portion of hardware demand.

AI infrastructure investment must ultimately face the test of commercial returns. If companies pour massive amounts of capital into building data centers but fail to generate sufficient revenue through subscriptions, advertising, cloud services, or enterprise software, the growth rate of capital expenditures may slow. The extent to which different chip companies benefit will also diverge, and not all AI-concept stocks will grow in tandem.

Kimi K3 could also intensify price competition in the model market. As open-source or open-weight models increasingly close the gap with the capabilities of closed-source models, the API profit margins of AI companies may come under pressure. To lower costs, model developers will place greater emphasis on custom chips, quantization technologies, inference optimization, and multi-vendor strategies, which could weaken the bargaining power of general-purpose GPU makers in certain markets.

Therefore, the impact of low-cost models on the AI chip industry is not a one-way positive. While it expands the AI application market on one hand, it also forces infrastructure companies to improve energy efficiency and reduce unit computing costs on the other.

This content was translated using AI and reviewed for clarity. It is for informational purposes only.

View Original
Disclaimer: The content of this article solely represents the author's personal opinions and does not reflect the official stance of Tradingkey. It should not be considered as investment advice. The article is intended for reference purposes only, and readers should not base any investment decisions solely on its content. Tradingkey bears no responsibility for any trading outcomes resulting from reliance on this article. Furthermore, Tradingkey cannot guarantee the accuracy of the article's content. Before making any investment decisions, it is advisable to consult an independent financial advisor to fully understand the associated risks.

Comments (0)

Click the $ button, enter the symbol, and select to link a stock, ETF, or other ticker.

0/500
Commenting Guidelines
Loading...

Recommended Articles

tradingkey.logo
Risk Warning: Our Website and Mobile App provides only general information on certain investment products. Finsights does not provide, and the provision of such information must not be construed as Finsights providing, financial advice or recommendation for any investment product.
Investment products are subject to significant investment risks, including the possible loss of the principal amount invested and may not be suitable for everyone. Past performance of investment products is not indicative of their future performance.
Finsights may allow third party advertisers or affiliates to place or deliver advertisements on our Website or Mobile App or any part thereof and may be compensated by them based on your interaction with the advertisements.
© Copyright: FINSIGHTS MEDIA PTE. LTD. All Rights Reserved.