Google Commentary: Gemini 3.8 Implements Cost Leadership Strategy
Google released Gemini 3.8 Flash on September 2, emphasizing high cost-performance, strong reasoning, and competitive pricing. Google’s strategic shift prioritizes low inference costs over peak model capability, leveraging its massive distribution network across billions of active users and proprietary TPU hardware optimizations. This low-cost advantage excels in high-frequency, low-complexity tasks. However, the primary risk involves potential trade-offs between aggressive cost reduction and maintaining long-term frontier model competitiveness amidst industry-wide capability homogenization.

Google Released Gemini 3.8 Yesterday With Extremely High Cost-Performance
On September 2, Google released Gemini 3.8 Flash. This marks the third release in the Flash series over the past six weeks, with a focus on enhancing reasoning, coding, and agent capabilities. Notably, by contrast, the flagship Pro series has not been updated in roughly six and a half months since the release of Gemini 3.1 Pro on February 19.
In terms of model capability evaluations, 3.8 Flash is already remarkably close to high-end models. On Artificial Analysis's Intelligence Index, 3.8 Flash High scored 59, up 3 points from 3.7 Flash High's 56, putting it on par with GPT-5.6 Sol xhigh and Grok 4.6 medium. In Google's official tests, 3.8 Flash achieved an HLE-Verified score of 54.9%, while also significantly outperforming the previous generation in software engineering and agent benchmarks.
The pricing impact is even more striking. The API price for 3.8 Flash is $0.75 per million input tokens and $3.75 per million output tokens. That means for the same overall capability score on Artificial Analysis, its price is just 19% of GPT-5.6 Sol and 15% of Claude Opus 5. Of course, scores do not fully equate to capability, but its cost advantage is undeniable.
My Assessment: Google’s Large Model Strategy Prioritizes Low Cost, High-Quality Models Are No Longer the Top Priority
Over the past few years, the central question in the competition among large AI models was who could build the most powerful model.
However, this competitive logic is shifting.
The reason is that as model capabilities continue to improve, an increasing number of real-world tasks have crossed the threshold of 'whether the model can complete them.' For highly complex scientific research, coding, and long-horizon agents, model capability still holds immense value; however, for the far more numerous low-complexity tasks, improving a model's score from 90 to 95 may yield far less actual value than reducing costs from 1 to 0.3.
Combined with three updates to the Flash model over six weeks and the temporary quiet surrounding the Pro model, this indicates that Google's large model strategy is becoming increasingly clear: prioritizing low costs.
Of course, you might notice that 3.8 is actually somewhat more expensive than 3.7. While the API unit prices for both model generations are identical, average output tokens grew by about 30%, and invocation turns for agent tasks also increased. As measured by Artificial Analysis, the per-task cost rose from approximately $0.40 to $0.58, an increase of about 45%. We believe Google is currently exploring the boundary between cost and performance along the low-cost route, but the low-cost strategy itself remains sound.
On the other hand, executive comments also corroborate this. In the second quarter of this year, Google CEO Sundar Pichai stated that demand for the Flash series is very strong because it sits at the sweet spot of performance and cost.
Clearly, Google is shifting its top-priority objective from:
"how powerful the strongest model can be made"
to:
"how low the cost of completing tasks can be driven."
There is a major difference between these two approaches. The latter model aligns better with the economics of future large-scale AI applications, as the vast majority of compute will ultimately not come from a tiny fraction of the most difficult problems, but rather from a massive volume of ordinary requests.
Why This Strategy Is Particularly Suited for Google
1. Google's distribution advantage brings in a vast user base, while a high volume of requests itself amplifies the value of low costs.
A fundamental difference between Google and other large model companies is that Google already has a user base.
In the second quarter of this year, monthly active users for AI Mode in Google Search surpassed 1 billion; on August 11, the Gemini App exceeded 1 billion monthly active users. An even larger distribution advantage stems from the entire Google product matrix: Gemini is now integrated into 13 Google products with user bases exceeding 1 billion each, 5 of which have over 3 billion users.
The developer side has also reached massive scale. Every month, over 9 million developers use Google's models, and its first-party model API currently processes approximately 22 billion tokens per minute, up from 16 billion just a quarter ago.
Therefore, the economic problem Google faces is different from that of an AI company acquiring users from scratch.
What it truly needs to solve is:
How to deliver AI to its existing billions of users at a sufficiently low cost.
At this point, identifying task complexity becomes extremely important.
For high-complexity tasks, customers primarily purchase capability. For scientific research or complex agent tasks, if a more powerful model can significantly improve the success rate, spending a few extra dollars per task is completely acceptable.
However, Google primarily faces low-complexity, high-frequency requests on a daily basis.
For instance, with general Q&A and simple agents, as long as both models can solve the problem correctly, increasing model capability by a few more percentage points yields very low marginal value to users. However, if the cost per inference drops from $0.50 to $0.20, multiplied by billions of users and constantly growing query frequencies, it creates a massive cost difference for LLM companies.
Therefore, a multiplier effect exists between distribution advantages and low-cost advantages.
The more users Google has, the greater the value generated by cost reductions per inference; the lower the cost, the more Google can further embed AI into existing products such as Search, Android, Workspace, and YouTube, thereby continuing to expand query volume.
This is also why low-cost models may hold higher strategic value for Google than for other LLM companies.
2. The low-cost advantage of TPUs will become the core deciding factor in low-complexity scenarios
For high-complexity tasks, cost differences are usually not the deciding factor.
If one model can accomplish scientific research, coding, or complex agent tasks that another model cannot, users primarily purchase the result; a slightly more expensive model will not change their choice.
However, the decision criteria for low-complexity tasks are completely different.
Once a task can be reliably completed by various models, the marginal returns from further improvements in model capability drop rapidly, making completion cost the core competitive factor.
This is precisely where TPUs (Google's proprietary chips) can maximize their value.
Google can simultaneously control chips, models, and inference infrastructure. The latest TPU 8i itself is designed specifically for inference workloads, and data released by Google shows an 80% increase in inference performance per dollar compared to the previous-generation TPU.
On the actual product side, these cost reductions have already begun to materialize. Alphabet disclosed in Q2 that even as AI Mode continuously incorporates more advanced models and capabilities, joint software-hardware optimization has driven the cost per AI Mode answer down to its lowest level since the product's launch.
TPUs endow Google with a key capability:
The capability to sustain a cost advantage amid the race for model capabilities.
In top-tier scenarios, this advantage might only marginally affect profit margins without capturing market share; however, in low-complexity, billion-user, high-frequency scenarios, cost itself dictates which AI features can be delivered at scale and who can ultimately sustain the largest volume of queries.
Biggest Risk: After Model Homogenization, Google Must Be One of the Surviving Players
For Google's strategy to truly hold, two conditions must be met simultaneously:
Frontier models gradually become homogeneous, while Google itself consistently remains in the frontier model cohort.
I believe the likelihood of these two conditions occurring simultaneously is rising over the next two to three years.
The probability of the first condition is extremely high. Model competition over the past few years has exhibited a distinct feature: it is becoming increasingly difficult for any company to sustain a lead for long. New training methods, inference techniques, and model architectures diffuse rapidly; after a company achieves a clear lead, other top-tier players usually catch up in subsequent rounds of model updates.
Therefore, the landscape most likely to emerge ultimately is one where a small number of companies remain within a homogeneous capability range over the long term.
However, under a low-cost strategy, whether the probability of Google consistently remaining in this cohort declines is precisely the biggest risk of this strategy. Over the past six weeks, Google has updated its Flash series three consecutive times; in contrast, the Pro series has gone without an update for six and a half months. If R&D resources and product goals continue to tilt toward cost efficiency, Google may sacrifice a portion of its frontier capabilities in exchange for lower inference costs.
This creates an inherent contradiction: Google's low-cost strategy delivers maximum value only when its own models remain in the top tier, but prioritizing low cost may weaken its ability to stay in that top tier.
For now, we are adopting a wait-and-see approach on this issue, as history repeatedly shows that in industries lacking a first-mover advantage, relentlessly giving chase may be inferior to waiting for the sector to reach a bottleneck and then catching up at full speed—which may well be Google's true playbook.
This content was translated using AI and reviewed for clarity. It is for informational purposes only.
Recommended Articles












Comments (0)
Click the $ button, enter the symbol, and select to link a stock, ETF, or other ticker.