World of the Developers
All calculators

Gemini API Cost Calculator

Calculate Google Gemini API costs across Flash, Flash-Lite, and Pro tiers for your exact token usage.

Usage

$1.50/M in · $0.15/M cached in · $9.00/M out · 1M context

Estimated cost

$180.00/month

Per request

$0.006

Per day

$6.00

Per year

$2,190.00

Compare across models

ModelProviderPer month
OpenAI$7.50
OpenAI$13.50
OpenAI$24.00
OpenAI$36.00
OpenAI$37.50
OpenAI$37.50
Google$46.50
Google$78.75
Anthropic$105.00
OpenAI$180.00
OpenAI$180.00
Google$180.00
OpenAI$187.50
Anthropic$210.00
OpenAI$225.00
OpenAI$240.00
Google$240.00
Anthropic$315.00
OpenAI$420.00
Anthropic$525.00
Anthropic$525.00
Anthropic$525.00
Anthropic$1,050.00

How this is calculated

Providers bill input and output tokens separately, per million tokens (MTok). Cached input tokens — a repeated prompt prefix the provider recognizes — are billed at a discounted rate.

cost = (input_tokens − cached_tokens) × input_rate
     + cached_tokens × cached_rate
     + output_tokens × output_rate

Pricing is verified against provider pricing pages as of 2026-09-01. AI and cloud pricing changes frequently — confirm the current rate on the provider's own pricing page before budgeting.

Frequently asked questions

What's the cheapest Gemini model?

Gemini 3.5 Flash-Lite is the cheapest current-generation model at $0.30/$2.50 per million input/output tokens, suited to high-volume, simple tasks.

Why does Gemini 3.1 Pro cost more above 200K tokens?

Google prices long-context requests on a separate, higher tier — Gemini 3.1 Pro rises from $2.00/$12.00 to $4.00/$18.00 per million input/output tokens once a request exceeds 200K tokens.

Is the Gemini 3.7 Flash price permanent?

No — it's an introductory rate. Google's published pricing shows Gemini 3.7 Flash at $0.75/$3.75 per million tokens through the end of 2026, doubling to $1.50/$7.50 starting January 1, 2027.

Does Gemini support prompt caching?

Yes — cached input tokens are billed at a fraction of the standard input rate (around 10%), the same mechanism OpenAI and Anthropic use for repeated prompt prefixes.