Gemini Flash

Family

Google · Gemini

Google's fast Gemini line, led by Gemini 3.5 Flash plus Live, TTS, Translate, and efficient variants.

fast efficient multimodal long-context cost-effective model-family
Updated July 6, 2026

Overview

This is a model family overview. For version-specific details, see the individual model entries linked below.

Gemini Flash is Google’s speed-and-cost tier, designed for tasks where throughput, latency, and price matter alongside strong multimodal reasoning. The family now spans stable production models, preview fast-model experiments, Live/TTS/Translate variants, and adjacent browser-control or agent surfaces. The center of gravity moved in May 2026 with Gemini 3.5 Flash, Google’s stable fast frontier model for coding, multimodal understanding, and long-horizon agent workflows.

Current Latest

Gemini 3.5 Flash is the current stable fast/agentic route in the Gemini API. Google’s current catalog also keeps Gemini 3 Flash as a preview entry with Gemini 3.5 Flash as the recommended replacement, Gemini 3.1 Flash-Lite as the stable efficient route, Gemini 3.1 Flash Live/TTS as adjacent preview audio surfaces, and Gemini 3.5 Live Translate as a narrower realtime translation route. Older Gemini 2.5 Flash entries remain useful compatibility and cost baselines, but Google’s deprecation page now lists October 16, 2026 shutdown dates for the stable 2.5 Flash and Flash-Lite endpoints.

Strengths

  • Very fast inference for latency-sensitive applications
  • Stable Gemini 3.5 Flash route for agentic coding and long-horizon workflows
  • Competitive pricing relative to Pro tiers
  • Full multimodal support across text, image, video, audio, and PDFs
  • 1M-token context windows on stable Flash and Flash-Lite
  • Stable Flash-Lite variant for the most cost-sensitive high-volume workloads
  • Preview-tier Gemini 3 Flash and Live/TTS/Translate variants for teams tracking newer fast-model direction

When to Choose Gemini Flash

  • High-volume processing where cost per request matters
  • Real-time applications requiring low latency
  • Bulk document analysis and extraction pipelines
  • Development prototyping before escalating to Pro or managed-agent routes
  • Applications where multimodal support is needed at scale
  • Teams that want a stable 3.5 Flash production lane while evaluating preview and media-specific variants

Access

  • Google AI Studio
  • Vertex AI and Gemini Enterprise Agent Platform deployment paths
  • Google Gemini consumer products
  • Third-party integrations via API