Website icon Xpert.Digital

Where is Gemini 3.5 Pro? Google's Gemini offensive: When speed becomes more important than ingenuity

Where is Gemini 3.5 Pro? Google's Gemini offensive: When speed becomes more important than ingenuity

Where is the Gemini 3.5 Pro? Google's Gemini offensive: When speed becomes more important than ingenuity – Image: Xpert.Digital

Price drop in AI: How Google is revolutionizing the market with Gemini 3.6 Flash

Cheaper instead of brilliant: Google's surprising change of strategy at Gemini

Google's new AI plan: Is the company abandoning the race for the smartest AI?

Google is drastically changing the rules of the game in the global AI race. Instead of engaging in an isolated battle with competitors like OpenAI and Anthropic for the most intelligent—and therefore most expensive—flagship model, the tech giant is focusing on radical efficiency. With its new model family centered around Gemini 3.6 Flash, the extremely power-efficient Flash-Lite, and the highly specialized but currently under wraps Cyber ​​model, a new metric is taking center stage: cost reduction with maximum practicality. While the long-awaited flagship Gemini 3.5 Pro remains unavailable, Google is repositioning itself as an indispensable foundational provider for the rapidly growing agent economy. But what does this strategic shift mean for the future of generative AI? And why does Google consider its latest security model too risky for the general public? The following article analyzes Google's pragmatic path to the next generation of AI and examines its massive impact on businesses and developers.

Why Google prefers to deliver three small models rather than fulfill one big promise

Google's latest move in the race among major AI providers follows a pattern that is increasingly becoming the norm in the industry: Instead of a single, monolithic advancement, the company is delivering an entire family of models tailored to different use cases. With Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, Google is launching the next evolutionary stage within just a few months of introducing Gemini 3.5 Flash in May. What's remarkable here is less the technical substance of the individual innovations than the underlying strategic message: Google is no longer primarily positioning itself as the provider of the most intelligent model, but rather as the provider of the most economically efficient AI ecosystem.

Frugality as a new status symbol in the AI ​​race

The central selling point surrounding Gemini 3.6 Flash is not increased intelligence, but cost reduction. Google advertises a roughly 17 percent reduction in token consumption compared to its predecessor, 3.5 Flash, based on measurements from the Artificial Analysis Index. In some benchmarks, such as DeepSWE, the cost per generated output token is said to decrease by as much as 65 percent. At the same time, the nominal price has also been reduced: While Gemini 3.5 Flash was billed at $1.50 per million input tokens and $9 per million output tokens, Gemini 3.6 Flash costs only $7.50 per million output tokens with unchanged input costs. This combination of lower consumption and a lower price per unit results in a clearly noticeable effect on the operating costs of companies that deploy generative AI at scale.

This focus on efficiency is no coincidence, but rather an expression of an economic reality currently shaping the entire industry. Large language models incur computational costs with each individual query, which add up to considerable sums when dealing with millions or billions of daily queries. Cloud providers and platform operators who want to scale agent systems that independently handle multi-stage tasks must drastically reduce token consumption per step; otherwise, the economic benefits of such systems will be consumed by infrastructure costs. Google addresses precisely this problem by offering a new model that, according to its own statements, requires fewer thought processes and tool calls to complete multi-stage workflows.

Technological advances beyond the mere question of cost

In addition to increased efficiency, Gemini 3.6 Flash also shows gains in classic performance benchmarks. In programming tasks, as measured by the DeepSWE test, the success rate rises from 37 to 49 percent. In the ML Research Benchmark MLE Bench, the model improves from 49.7 to 63.9 percent. Progress is also evident in knowledge-based tasks: The GDPval-AA score climbs from 1349 to 1421 points. Particularly noteworthy is the improvement in computer literacy, specifically the AI's ability to independently navigate graphical user interfaces: The OSWorld-Verified score increases from 78.4 to 83 percent. Furthermore, the model's knowledge base has been updated from January 2025 to March 2026, a significant practical advantage in today's rapidly evolving technological landscape.

Despite these advances, Gemini 3.6 Flash remains clearly positioned as a mid-range model. In direct comparison tests on platforms like Arena.ai, it ranks behind top-of-the-line models such as Claude Opus or the new GPT 5.6-Sol. This isn't a weakness, but a deliberate design choice: Google is foregoing the claim of offering the most powerful model overall and instead focusing on a segment where speed, cost, and sufficient quality are more important than the theoretical maximum in terms of intelligence. This strategy is reminiscent of classic market segmentation in other technology sectors, such as processors, where manufacturers consistently offer efficient mid-range products for mass-market applications alongside high-performance chips.

The new base layer for agent economics

With Gemini 3.5 Flash-Lite, Google is pursuing an even more radical efficiency goal. This model is explicitly designed as the foundation for building and scaling AI agents—systems that can independently process complex tasks in multiple steps without requiring human intervention at each stage. It significantly outperforms its predecessor, 3.1 Flash-Lite, in benchmarks for both knowledge-based and programming tasks, while operating costs remain exceptionally low at $0.30 per million input tokens and $2.50 per million output tokens. The model's flexibility is remarkable: depending on the use case, it can be configured to operate particularly quickly and cost-effectively or to rely more heavily on its reasoning capabilities. The latter becomes necessary when tasks need to be processed collaboratively by several specialized subroutines, known as subagents.

This architectural decision reveals the direction the entire industry is heading. Many market observers believe the future of AI lies not in ever-larger individual models, but in networked systems of smaller, specialized models that, working together, map complex processes. A single large calculation by a powerful model is replaced by many small, inexpensive queries to lean models that collaborate in parallel or sequentially. A model like Flash-Lite is ideally suited to this architecture because it reduces the cost per query to such an extent that a high number of queries remains economically viable. The announcement that Flash-Lite will also deliver AI results in Google web search further underscores the central role Google assigns to this lean model in the mass market.

A tool with two faces: Cybersecurity as a business field

The third new release, Gemini 3.5 Flash Cyber, marks a significant expansion of the Flash series into a highly sensitive area. This model has been specifically trained to find and analyze security vulnerabilities in software code and to suggest remediation measures. This specialization is based on the already efficient Gemini 3.5 Flash, which has been further refined through targeted training to analyze large codebases. Instead of sending a single, expensive query to a very large model, the overarching security agent Codemender calls the streamlined Cyber ​​model multiple times to examine different execution paths of a program in parallel. The individual partial results are then consolidated into a comprehensive report.

The performance of this approach can be seen in concrete figures. In a test on the complex JavaScript engine V8, the cyber model identified 55 distinct, confirmed vulnerabilities, while the general Flash model found only 47 and Anthropic's Claude Opus 4.6 a mere 36. Ten of the discovered issues were found exclusively by the specialized cyber model. In the CyberGym benchmark, which evaluates AI agents against hundreds of real-world software vulnerabilities, the model, in combination with multiple runs, achieves a level of performance comparable to significantly larger and more expensive security models from vendors like Anthropic and OpenAI. In an internal test by Google's cloud security team, the model discovered several critical vulnerabilities in production systems within just two hours, including a memory corruption vulnerability for which it also generated a working exploit that bypassed common protection mechanisms such as address space layout randomization.

This enormous ability to identify vulnerabilities in code is also the model's greatest risk. Google itself openly acknowledges the technology's dual-use nature: the same capability that allows defenders to close vulnerabilities could be used by attackers to exploit those very same gaps before they are patched. For this reason, Google has decided not to make the model publicly available for the time being, but rather to share it exclusively through its own security tool, Codemender, as part of a tightly controlled pilot program with governments and selected, trusted partners. No concrete timeline for expanding access has been given. This reluctance reflects a growing awareness across the industry that highly specialized AI systems in the security field cannot be released under the same rules as general chatbot applications.

 

A new dimension of digital transformation with 'Managed AI' (Artificial Intelligence) - Platform & B2B solution | Xpert Consulting

A new dimension of digital transformation with 'Managed AI' (Artificial Intelligence) – Platform & B2B solution | Xpert Consulting - Image: Xpert.Digital

Here you will learn how your company can implement customized AI solutions quickly, securely and without high entry barriers.

A managed AI platform is your all-inclusive, worry-free solution for artificial intelligence. Instead of dealing with complex technology, expensive infrastructure, and lengthy development processes, you receive a ready-made solution tailored to your needs from a specialized partner – often within just a few days.

The key advantages at a glance:

⚡ Rapid implementation: From idea to ready-to-use application in days, not months. We deliver practical solutions that create immediate added value.

🔒 Maximum data security: Your sensitive data stays with you. We guarantee secure and compliant processing without sharing data with third parties.

💸 No financial risk: You only pay for results. High upfront investments in hardware, software, or personnel are completely eliminated.

🎯 Focus on your core business: Concentrate on what you do best. We take care of the entire technical implementation, operation, and maintenance of your AI solution.

📈 Future-proof & scalable: Your AI grows with you. We ensure continuous optimization and scalability, and flexibly adapt the models to new requirements.

More information here:

 

Google postpones Gemini 3.5 Pro and focuses on a major Flash offensive

The mystery surrounding the missing flagship model

What's striking about the entire announcement is what's missing: Gemini 3.5 Pro, the most powerful model of the current generation, was already expected at Google I/O in May and hasn't materialized since. Google merely explains that the new version is currently being tested with select partners and will be made available after successful completion of these tests. The company still hasn't given a specific release date, but emphasizes that the wait shouldn't be much longer. From the perspective of several industry observers, this repeated delay suggests that the Pro version hasn't yet achieved the desired results internally in complex programming and reasoning tasks, which is why Google apparently prefers to wait rather than release a product that doesn't meet its own quality standards or the expectations of its competitors.

This reticence sheds an interesting light on the competitive dynamics among leading AI providers. While OpenAI and Anthropic are releasing new top-of-the-line models in rapid succession, Google is currently visibly focusing on the mid-range and entry-level performance segments. This can be interpreted as a sign of caution, but also as shrewd market positioning: Instead of presenting a potentially still-developing flagship model, Google is securing market share across the board with affordable, fast, and practical models that already reliably cover the vast majority of everyday business applications. For many business use cases, such as document processing, reporting, or simple programming tasks, the absolute peak performance of a model is of secondary importance anyway, as long as speed and cost are right.

Availability and integration into the existing ecosystem

The practical adoption of the new models is remarkably rapid and widespread. Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are already available in Google's AI Studio, Android Studio, and the Gemini Enterprise Agent Platform. Additionally, 3.6 Flash has been integrated into the Antigravity development environment and the Gemini Enterprise app, while both models are also available in the regular Gemini end-user app. Particularly noteworthy is the announcement that Flash-Lite will soon also deliver the AI-generated results of Google web searches, which, given the enormous volume of queries from this search engine, underscores the importance of a cost-effective model for Google in a completely new dimension. The rapid adoption is also evident outside of Google's own ecosystem: Gemini 3.6 Flash was integrated into GitHub Copilot on the day of its announcement, highlighting the increasing integration of major AI providers with established developer tools.

This broad and simultaneous integration across diverse product lines demonstrates Google's systematic approach to seamlessly rolling out new model generations across its entire product portfolio. Unlike previous model changes, which often involved noticeable gaps between announcement and actual availability in various applications, this rollout appears to be a meticulously orchestrated process. This suggests a mature internal infrastructure that allows for the near-simultaneous deployment of a new model across all touchpoints with developers and end users.

Looking ahead: The announcement of Gemini 4

In addition to the three specific model releases, Google's announcement also includes a glimpse into the next major generation. The company confirms that the so-called pre-training, the fundamental initial training phase for Gemini 4, has already begun. Google does not provide specific technical details regarding the architecture, size, or planned capabilities of the upcoming large-scale model, but emphasizes its excitement about the progress made so far. This deliberately vague wording is typical of the communication strategy of large AI labs, which on the one hand raise expectations, but on the other hand do not want to make verifiable promises that might later prove premature.

Considering the release cycle of recent years, in which Google typically introduced new base models at its I/O developer conference in the spring and followed up with major generational leaps in the fall or winter, a Gemini 4 release toward the end of this year seems plausible. For companies planning their AI strategy for the long term, this means that anyone investing in the Flash family today does so with the awareness that a new technological foundation could follow in just a few months, which in turn would require new migration decisions. This short half-life of AI model generations presents IT departments and software architects with the ongoing challenge of designing their systems in such a way that a change in the underlying models can be implemented as smoothly as possible.

Context: What this model strategy reveals about Google's market position

Overall, the recent announcement demonstrates that the competition among leading AI companies is increasingly splitting into two parallel dimensions. On the one hand, there is the race for the most intelligent, most powerful model, a field in which Anthropic and OpenAI currently dominate the headlines with their flagship products, and in which Google is visibly lagging behind with its delayed Pro version. On the other hand, a second, potentially more economically significant battleground is emerging: the competition for the most efficient, cost-effective, and scalable infrastructure for the everyday, mass deployment of AI in businesses, search engines, and agent systems. In this second arena, Google is very deliberately and consistently positioning itself at the forefront with its Flash family of products.

This dual strategy is economically sound. For the vast majority of practical applications in businesses, from automated document processing and customer service chatbots to simple programming tasks, the absolute top-tier intelligence of a model is not the deciding factor, but rather the balance between cost, speed, and sufficient quality. Cloud providers who have to process millions of requests per day benefit more from a double-digit percentage reduction in token costs than from a marginal improvement in the reasoning capabilities of an expensive top-of-the-line model. Google appears to be consistently translating this insight into its product strategy, simultaneously sending a signal to developers and enterprise customers: Anyone wanting to build scalable AI agents or cost-effective search integrations today will find a solid, economically viable foundation in the Flash family, regardless of when the large Pro model is ultimately released.

At the same time, the introduction of the specialized cyber model reveals a new facet of the AI ​​competitive landscape that extends beyond mere efficiency considerations: the increasing use of AI as an active tool in cybersecurity, with all the associated opportunities for defenders and risks in the event of misuse. The controlled, phased release of this model demonstrates that Google is aware of its social responsibility in releasing particularly powerful security tools and is prepared to limit commercial reach, at least temporarily, in favor of security concerns. How this balance between broad availability and controlled access evolves in the future is likely to become one of the most interesting areas to observe as AI models become increasingly capable of independently analyzing and manipulating complex technical systems.

 

Your global marketing and business development partner

☑️ Our business language is English or German

☑️ NEW: Correspondence in your native language!

 

Konrad Wolfenstein

I and my team are happy to be available to you as your personal advisor.

You can contact me by filling out the contact form here wolfenstein@xpert.digital:or simply call me at +49 7348 4088 965. My email address is

I'm looking forward to our joint project.

 

 

☑️ SME support in strategy, consulting, planning and implementation

☑️ Creation or realignment of the digital strategy and digitization

☑️ Expansion and optimization of international sales processes

☑️ Global & Digital B2B trading platforms

☑️ Pioneer Business Development / Marketing / PR / Trade Fairs

 

🎯🎯🎯 Data-driven B2B industry hub as a quasi-in-house solution

The quasi-in-house solution: How Xpert.Digital closes operational gaps in B2B marketing and sales – Smart Content-Driven Business - Image: Xpert.Digital

Xpert.Digital is a data-driven B2B industry hub led by Konrad Wolfenstein . The company acts as an external, quasi-in-house solution for industrial partners, closing operational gaps in marketing, content, and sales – without requiring additional resources on the client side.

More information here:

Leave the mobile version