gemini
Gemini is the multimodal large language model developed by Google DeepMind, supporting the understanding and generation of text, images, audio, and video, with native long context of 1M–2M tokens. It offers three tiers—Pro / Flash / Flash-Lite—and supports image generation (Imagen / Gemini Image).
1M–2M token context window; implicit cache (cache_read 25% of input); explicit cache (with configurable TTL); visual understanding (images/video/PDF); audio input; image generation (Gemini 3 Pro Image / Flash Image); Function Calling; code execution; Google Search tool; JSON mode.
Multimodal Q&A; video content understanding and retrieval; image generation and editing; long-document analysis; code generation; agent workflows; Google Workspace integration.
Image generation is billed per image; video generation is billed per second; no on-premises deployment (except via Vertex AI); the enterprise edition requires a Google Cloud account; some models require requesting preview access.
| Model | Input / 1M | Output / 1M | Cache Read / 1M | Per 1M |
|---|---|---|---|---|
| 1M | ||||
| 1M | ||||
| 1M | ||||
| — | — | 1M | ||
| - | 1M | |||
Gemini 3.1 Pro gemini-3.1-pro |
| $2/M$1.4/M-30% |
| $12/M$8.4/M-30% |
| $0.2/M$0.14/M-30% |
| 1M |
Gemini 3.1 Flash Lite gemini-3.1-flash-lite | $0.25/M$0.18/M-30% | $1.5/M$1.05/M-30% | $0.025/M$0.0175/M-30% | 1M |
Gemini 3 Flash gemini-3-flash | $0.5/M$0.35/M-30% | $3/M$2.1/M-30% | $0.05/M$0.035/M-30% | 1M |
Gemini 3 Pro Image gemini-3-pro-image | $2/image$1.4/image-30% | $12/image$8.4/image-30% | - | 1M |
Gemini Embedding 2 gemini-embedding-2 | $0.2/M$0.14/M-30% | $0/M | - | 1M |
Gemini 2.5 Flash Lite gemini-2.5-flash-lite | $0.08/M$0.06/M-30% | $0.3/M$0.21/M-30% | $0.0075/M$0.0053/M-30% | 1M |