Quick Facts
- Gemini 3.8 Flash scored 73.7% on the DeepSWE v1.1 software engineering benchmark, topping GPT-5.6 Sol and falling just below Claude Opus 5’s 74.0%.
- A companion model, Gemini 3.8 Flash Cyber, scored 86.2% on CyberGym and produced 2.6 times more correct vulnerability patches in Chrome than the best available commercial models.
- Introductory pricing runs $0.75 per million input tokens through December 31, 2026, rising to $1.50 per million on January 1, 2027.
Google launched two new AI models on September 2, 2026: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The release came three weeks after Gemini 3.7 Flash and marks the third Flash update in three months.
The general-purpose model, known internally as Skimaki, targets software engineering, multi-step reasoning, and long-running agentic tasks. On the Terminal-Bench 2.1 benchmark, which measures how reliably a model completes real command-line and coding tasks end to end, 3.8 Flash scored 90.8% compared to 81.6% for 3.7 Flash. On HLE-Verified, a test covering multi-step STEM and professional reasoning, it scored 54.9%.
Google executives Tulsee Doshi and Raluca Ada Popa wrote in a blog post: “These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively.”
Speed and Cost
Gemini 3.8 Flash generates output at 304.6 tokens per second through Google’s API. That compares to a median of 70.8 tokens per second for reasoning models in a similar price tier. The model supports a 1-million token context window with up to 64,000 output tokens.
The model accepts text, image, video, audio, and PDF inputs. Reasoning effort is tunable at low, medium, and high settings, matching the 3.7 Flash API structure.
The added capability comes at a cost. Gemini 3.8 Flash runs about 40% more expensive per task than its predecessor because it outputs more tokens and executes more turns when acting as an agent. Developers should factor that overhead into any agentic pipeline cost estimates.
A Dedicated Security Model
Gemini 3.8 Flash Cyber is a separate, security-specialized variant. It scored 86.2% on the CyberGym benchmark, which tests a model’s ability to identify vulnerabilities in C and C++ code. Google’s internal evaluation found the model achieved a success rate above 70% when discovering vulnerabilities across codebases spanning 20 programming languages.
Security firm Wiz found that Gemini 3.8 Flash Cyber achieved 7.5 to 9.7 percentage points higher recall on an internal penetration testing benchmark at 2.3 to 5.2 times lower cost compared to other leading frontier models.
Access to the Cyber model is restricted. Google is distributing it through the Fairwind Program, which initially covers government agencies, Google Cloud customers, and cybersecurity partners, with priority given to critical infrastructure operators and maintainers of widely used open-source software.
Fairwind and CodeMender
The Fairwind Program pairs Gemini 3.8 Flash Cyber with CodeMender, Google’s AI agent for vulnerability remediation. The combination is designed to generate verified, deployment-ready patches inside an organization’s secure cloud environment in minutes rather than weeks.
Doshi said: “We’re really excited about being able to provide an offering to defenders that is a fraction of the cost, much faster, while still showcasing that frontier-level performance.”
Where the Model Is Live
Gemini 3.8 Flash is available now in the Gemini app for Google AI Pro and Ultra subscribers, in AI Mode, in Gemini in Google Sheets, and through Google AI Studio and Android Studio. The security model remains limited to Fairwind Program participants.
Read more: Google launches two Gemini 3.8 models with cutting-edge reasoning capabilities
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
