
Google quietly launched the Gemini 3.8 Flash model today (the 2nd) on the Google DeepMind website. Based on Gemini 3.7 Flash, it delivers performance improvements in software engineering and agentic knowledge workflows, while continuing to support customizable effort levels to control the balance of quality, cost, and latency.
However, as of IT Home's publication time, Google had not officially announced the model's existence through a press release or social media platform.
Gemini 3.8 Flash is based on Gemini 3.7 Flash and features a maximum 1M-token context window, with support for up to 64K tokens of text output. Google also published evaluation results from multiple benchmarks covering coding, knowledge processing, multimodal capabilities, long-context processing, computer use, and scientific reasoning.
The results as of September 2026 are as follows:

According to the introduction, Gemini 3.8 Flash is designed for individual users, developers, and enterprises. It is suitable for deploying general-purpose agents at scale and at lower cost, with typical uses including software engineering, agent tasks, and complex knowledge-processing workflows.
Gemini 3.8 Flash also has some common limitations of foundation models, such as the possibility of producing hallucinations. Google is continuing to improve the model's resistance to jailbreak attacks and has recently further strengthened safeguards in frontier safety areas.
The model may occasionally respond slowly or time out. To achieve better performance, Gemini 3.8 Flash may consume more tokens at times, especially at higher reasoning intensities.
Gemini 3.8 Flash has a knowledge cutoff date of March 2026. Information in some areas may have been updated later, while knowledge in other areas may still date back to January 2025, consistent with the Gemini 3 model family.
References
Official Gemini 3.8 Flash introduction
