Gemini API Cost Calculator
Calculate Google Gemini API costs across Flash, Flash-Lite, and Pro tiers for your exact token usage.
Usage
$1.50/M in · $0.15/M cached in · $9.00/M out · 1M context
Estimated cost
$180.00/month
Per request
$0.006
Per day
$6.00
Per year
$2,190.00
Compare across models
| Model | Provider | Per month |
|---|---|---|
| OpenAI | $7.50 | |
| OpenAI | $13.50 | |
| OpenAI | $24.00 | |
| OpenAI | $36.00 | |
| OpenAI | $37.50 | |
| OpenAI | $37.50 | |
| $46.50 | ||
| $78.75 | ||
| Anthropic | $105.00 | |
| OpenAI | $180.00 | |
| OpenAI | $180.00 | |
| $180.00 | ||
| OpenAI | $187.50 | |
| Anthropic | $210.00 | |
| OpenAI | $225.00 | |
| OpenAI | $240.00 | |
| $240.00 | ||
| Anthropic | $315.00 | |
| OpenAI | $420.00 | |
| Anthropic | $525.00 | |
| Anthropic | $525.00 | |
| Anthropic | $525.00 | |
| Anthropic | $1,050.00 |
How this is calculated
Providers bill input and output tokens separately, per million tokens (MTok). Cached input tokens — a repeated prompt prefix the provider recognizes — are billed at a discounted rate.
cost = (input_tokens − cached_tokens) × input_rate
+ cached_tokens × cached_rate
+ output_tokens × output_ratePricing is verified against provider pricing pages as of 2026-09-01. AI and cloud pricing changes frequently — confirm the current rate on the provider's own pricing page before budgeting.
Frequently asked questions
What's the cheapest Gemini model?
Gemini 3.5 Flash-Lite is the cheapest current-generation model at $0.30/$2.50 per million input/output tokens, suited to high-volume, simple tasks.
Why does Gemini 3.1 Pro cost more above 200K tokens?
Google prices long-context requests on a separate, higher tier — Gemini 3.1 Pro rises from $2.00/$12.00 to $4.00/$18.00 per million input/output tokens once a request exceeds 200K tokens.
Is the Gemini 3.7 Flash price permanent?
No — it's an introductory rate. Google's published pricing shows Gemini 3.7 Flash at $0.75/$3.75 per million tokens through the end of 2026, doubling to $1.50/$7.50 starting January 1, 2027.
Does Gemini support prompt caching?
Yes — cached input tokens are billed at a fraction of the standard input rate (around 10%), the same mechanism OpenAI and Anthropic use for repeated prompt prefixes.