Updated 4 Sep 2026: Updates the Flash price/performance story from Gemini 3.7 to 3.8. Three weeks after 3.7, Google shipped 3.8 Flash at the same introductory token price with stronger coding/agent benchmarks, wider Google-product availability and a separate restricted 3.8 Flash Cyber model through Fairwind.

Key details

  1. Gemini 3.8 Flash launched September 2, 2026, three weeks after Gemini 3.7 Flash.
  2. Introductory API pricing remains $0.75/1M input tokens and $3.75/1M output tokens through December 31, 2026.
  3. Google says the published January 1, 2027 price remains $1.50/1M input and $7.50/1M output.
  4. Google reports gains across software engineering, long-horizon agents and multi-step reasoning while keeping Flash-tier speed.
  5. The model is available through the Gemini API, Google AI Studio, Antigravity, Gemini Enterprise and selected Google consumer products including AI Mode in Search.
  6. Gemini 3.8 Flash Cyber is a separate restricted-access model for defensive cybersecurity.
  7. The Fairwind Program limits Cyber access to vetted organizations and defensive/research uses.
  8. Independent reporting notes that higher reasoning effort can increase tokens consumed per task even at the unchanged list price.

What builders should take away

  1. Re-run your 3.7 evaluation set against 3.8 rather than assuming the version bump is automatically cheaper or better for your workload.
  2. Track cost per completed task, not just price per million tokens; include reasoning tokens, retries and tool-call loops.
  3. Keep the January 2027 standard-price step in your margin model even if 3.8 improves task success during the promotional period.
  4. If you use Gemini in Search-facing or agent workflows, test behavior separately across API and Google product surfaces rather than assuming one rollout behaves identically everywhere.
  5. Security teams considering Flash Cyber should treat Fairwind as a governed-access product with organizational eligibility and controls, not a model ID ordinary applications can freely route to.

What changed

On September 2, 2026, Google released Gemini 3.8 Flash only three weeks after Gemini 3.7 Flash. Google kept the same introductory Gemini API price of $0.75 per million input tokens and $3.75 per million output tokens through December 31, with the published January 1, 2027 rate still doubling to $1.50/$7.50. Google reports materially stronger software-engineering, agentic and multi-step reasoning performance while maintaining Flash-tier speed. The general model is available through the Gemini API, Google AI Studio, Antigravity, Gemini Enterprise and consumer surfaces including AI Mode in Google Search. Google simultaneously launched Gemini 3.8 Flash Cyber, a more permissive cybersecurity variant restricted to approved defenders through the new Fairwind Program.

Why it matters

For production agent systems, 3.8 changes the routing decision again without changing the headline token price. Builders can potentially buy more capability at the same list rate, but per-token pricing is not the same as per-task cost: more aggressive reasoning and longer trajectories can consume more output tokens, and early independent reporting points to that trade-off. The separate Cyber model also shows a new access pattern for frontier dual-use capability: general developers receive the standard model while a more capable security variant sits behind organizational vetting and use restrictions.

Google kept the list price while moving the capability baseline

Gemini 3.8 Flash launches at the same introductory $0.75/1M input and $3.75/1M output rate as 3.7 Flash. Google says the model materially improves long-horizon software engineering, agentic tasks and specialized reasoning, making the relevant production question total task cost and success rate rather than nominal token price alone.

More reasoning can still make a task cost more

The Verge reports early analysis indicating 3.8 can use materially more tokens on some tasks even though the per-token rate is unchanged. Teams should therefore compare end-to-end cost, retries, tool calls and wall-clock completion rather than assuming identical list pricing means identical workload economics.

The model is already spread across developer and consumer surfaces

Google says 3.8 Flash is available to developers through the Gemini API and AI Studio, to enterprises through Gemini Enterprise, and to Google AI Pro/Ultra users in the Gemini app, Google Sheets and AI Mode in Search. This makes the release both an API model change and a Search/consumer-model rollout.

Flash Cyber separates defensive capability from general access

Gemini 3.8 Flash Cyber shares the same foundation but has more permissive cybersecurity mitigations. Google limits it to approved Fairwind participants such as governments, critical-infrastructure operators, security partners and software maintainers, with governance requirements including controlled user access and defensive-use restrictions.

The January price step still matters

The introductory price remains scheduled to expire December 31. Google says standard pricing from January 1, 2027 will be $1.50/1M input and $7.50/1M output, so long-lived agent economics should still be modeled at both current and post-promotion rates.

What to watch next

  • Independent agent and coding benchmarks that measure total task cost as well as success rate.
  • Whether Google keeps shipping Flash revisions at the current unusually rapid cadence.
  • Any change to the December 31 promotional-pricing deadline.
  • Broader availability or changed eligibility for Gemini 3.8 Flash Cyber and Fairwind.
  • How Search AI Mode citation/link behavior changes as 3.8 Flash becomes part of the consumer Search stack.

Still unclear

  • Google’s strongest benchmark claims are project-authored and may not transfer to a specific production workload.
  • Early per-task cost estimates depend heavily on reasoning configuration and evaluation harness.
  • Gemini 3.8 Flash Cyber is restricted and cannot be assumed to have the same access, safeguards or deployment surface as the general model.
  • The rapid release cadence means 3.8 may itself be superseded quickly.

Sources

Direct reading behind this dossier.

3 sources

Discussion

Discussion is reader-contributed. Comments are not part of the BTN dossier or its editorial evidence.

0 visible comments

Join the discussion

Keep comments useful and relevant. Reader contributions may be moderated and are not BTN editorial evidence.

Sign in to comment