Gemini 1.5 Flash
Gemini 1.5 Flash was Google's speed-optimized multimodal large language model, released in 2024 and shut down on 29 September 2025 together with the rest of the Gemini 1.5 line, when Google directed developers to Gemini 2.0 Flash and Flash-Lite. It processed text, image, audio, and video input at lower latency and lower cost than the Pro variant. It was built for fast responses and a low cost per token while keeping the multimodal input types and the long context window of the 1.5 generation. Google's Gemini 1.5 technical report describes Flash as a transformer decoder model distilled online from the larger Gemini 1.5 Pro, which is itself a sparse mixture-of-experts model, and tuned for low latency and high throughput on TPUs; Google never described Flash itself as a mixture-of-experts model. It handled multimodal understanding, text generation, image analysis, conversational AI, and coding assistance, and was positioned for workloads that needed fast responses more than maximum quality.
While it was available, Gemini 1.5 Flash was used for real-time customer service systems, interactive applications, content moderation, automated analysis pipelines, and other high-volume workloads that needed multimodal processing at low cost. It was reached through the Gemini API and Vertex AI, and supported streaming as well as batch processing.
Related terms
Related services: Agentic AI consulting.
Ready to put agentic AI to work?
Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.