Gemini 4 Argon Pricing: API Cost, Token Calculator & 2027 Price Increase
Gemini 4 Argon Pricing: API Cost, Token Calculator & 2027 Price Increase
Google has announced unusually aggressive introductory pricing for its new frontier model, Gemini 4 Argon.
When Argon becomes available to paid API customers, Google says the introductory API price will be:
$2 per 1 million input tokens
and:
$10 per 1 million output tokens
Cached input tokens will receive a 95% discount from the normal input-token price.
But that introductory rate will not last indefinitely.
Google’s official launch materials state that the introductory pricing applies through December 31, 2026.
Beginning January 1, 2027, Gemini 4 Argon pricing is scheduled to increase to:
$4 per 1 million input tokens
and:
$20 per 1 million output tokens
In other words, the standard Argon token rates are scheduled to double after the introductory period.
That makes cost planning especially important for developers considering Argon for coding agents, enterprise research, cybersecurity, large document processing or other long-running AI workflows.
Here is how Gemini 4 Argon pricing works and what common API workloads could cost.
Gemini 4 Argon Pricing at a Glance
| Token type | Through Dec. 31, 2026 | From Jan. 1, 2027 |
|---|---|---|
| Input tokens | $2 / 1M | $4 / 1M |
| Output tokens | $10 / 1M | $20 / 1M |
| Cached input | 95% off input rate | Check current Google pricing |
| Price increase | — | 100% |
Google announced the pricing in its official Gemini 4 Argon launch post.
If you are looking for information about when you can actually use Argon, see our Gemini 4 Argon access and availability guide.
How Much Does Gemini 4 Argon Cost Per Token?
API providers typically publish prices per million tokens rather than per individual token.
For Gemini 4 Argon during the introductory period:
Input: $2 ÷ 1,000,000 = $0.000002 per token
Output: $10 ÷ 1,000,000 = $0.00001 per token
That means output tokens cost:
$10 ÷ $2 = 5× more than input tokens
The same 5-to-1 relationship remains under Google’s announced post-introductory pricing:
$20 output ÷ $4 input = 5×
That matters because applications producing very long responses can spend considerably more on output than input.
Gemini 4 Argon Cost Calculator Formula
The basic API-cost formula is straightforward.
Introductory pricing
Input cost:
Input tokens ÷ 1,000,000 × $2
Output cost:
Output tokens ÷ 1,000,000 × $10
Total:
Input cost + Output cost
Pricing from January 1, 2027
Input cost:
Input tokens ÷ 1,000,000 × $4
Output cost:
Output tokens ÷ 1,000,000 × $20
Total:
Input cost + Output cost
These calculations cover only the announced token rates.
A real application can involve other costs depending on the Google services, infrastructure, storage, tools or external APIs used around the model.
Example 1: 100,000 Input Tokens + 20,000 Output Tokens
Consider a relatively modest API request.
Introductory price
Input:
100,000 ÷ 1,000,000 × $2 = $0.20
Output:
20,000 ÷ 1,000,000 × $10 = $0.20
Total:
$0.20 + $0.20 = $0.40
So the simplified Argon token cost would be:
$0.40
From January 1, 2027
Input:
100,000 ÷ 1,000,000 × $4 = $0.40
Output:
20,000 ÷ 1,000,000 × $20 = $0.40
Total:
$0.80
The same workload would therefore double from approximately:
$0.40 → $0.80
Example 2: 1 Million Input Tokens + 100,000 Output Tokens
This could represent a large research or code-analysis workflow.
Introductory price
Input:
1M × $2 = $2
Output:
100,000 ÷ 1M × $10 = $1
Total:
$3
2027 price
Input:
1M × $4 = $4
Output:
100,000 ÷ 1M × $20 = $2
Total:
$6
So this simplified workload changes from:
$3 → $6
Example 3: 1 Million Input + 1 Million Output Tokens
Gemini 4 Argon is designed for unusually long, complex workflows, so large-output examples matter.
Introductory rate
Input:
$2
Output:
$10
Total:
$12
From January 1, 2027
Input:
$4
Output:
$20
Total:
$24
That is a difference of:
$12 per request
at this scale.
If an application ran that workload 1,000 times:
Introductory period
1,000 × $12 = $12,000
2027 rates
1,000 × $24 = $24,000
This illustrates why even small-looking per-million-token differences matter significantly at enterprise scale.
Example 4: Large AI Agent Workflow
Imagine a long-running coding or enterprise agent consuming:
5 million input tokens
and:
500,000 output tokens
Introductory pricing
Input:
5 × $2 = $10
Output:
0.5 × $10 = $5
Total:
$15
2027 pricing
Input:
5 × $4 = $20
Output:
0.5 × $20 = $10
Total:
$30
For 10,000 equivalent runs, the theoretical token cost would move from:
$150,000
to:
$300,000
This is why organizations planning high-volume Argon deployments should model costs using the post-introductory price rather than assuming the launch rate will continue indefinitely.
Gemini 4 Argon Price Calculator Table
Here are simplified estimates using Google’s announced rates.
| Input | Output | Intro Cost | 2027 Cost |
|---|---|---|---|
| 100K | 20K | $0.40 | $0.80 |
| 250K | 50K | $1.00 | $2.00 |
| 500K | 100K | $2.00 | $4.00 |
| 1M | 100K | $3.00 | $6.00 |
| 1M | 500K | $7.00 | $14.00 |
| 1M | 1M | $12.00 | $24.00 |
| 5M | 500K | $15.00 | $30.00 |
| 10M | 1M | $30.00 | $60.00 |
These estimates assume all tokens are billed at the standard announced input/output rates and do not include additional services.
How Does Gemini 4 Argon Cached Input Pricing Work?
Google says cached input tokens are priced at:
95% below the normal input-token price
during the introductory period.
A 95% discount means the user pays:
5% of the normal input rate
At the $2 introductory input price:
$2 × 5% = $0.10 per 1 million cached input tokens
That can create substantial savings when an application repeatedly sends the same large context.
For example, suppose a system repeatedly uses a one-million-token reference corpus.
Without caching at the introductory input rate:
$2 per request
With a 95% cached-input discount:
approximately $0.10 per million cached input tokens
The theoretical saving on that cached portion would be:
$2.00 − $0.10 = $1.90
or:
95%
Google’s exact caching rules, retention behavior and future pricing should be checked in the Gemini API documentation before designing a production cost model.
Why Context Caching Can Matter for Argon
Argon is designed for long-running professional workflows.
Potential workloads include:
- analyzing a large codebase;
- repeatedly consulting company documentation;
- working across legal files;
- financial research;
- cybersecurity analysis;
- autonomous software engineering; and
- multi-agent workflows.
Without caching, the same reference material might need to be processed repeatedly.
For example:
Large codebase → agent step 1 → agent step 2 → agent step 3
If the model has to reprocess the same context every time, costs can accumulate quickly.
Caching can make repeated-context workflows more economical when Google’s caching requirements are met.
Is Gemini 4 Argon Cheap?
That depends on what you compare it with and how you use it.
At its introductory rates, Argon is priced at:
$2 input / $10 output per million tokens
Google is positioning Argon as a frontier model rather than a low-cost lightweight model.
The important economic question is therefore not simply:
“Is $10 per million output tokens cheap?”
It is:
“Does Argon complete enough difficult work to justify its token cost?”
A more capable model can sometimes reduce overall project cost if it:
- solves the problem in fewer attempts;
- needs less human correction;
- automates more steps;
- produces better code;
- handles larger workflows; or
- replaces multiple calls to weaker models.
On the other hand, using a frontier model for easy classification or rewriting tasks may be unnecessarily expensive.
Gemini 4 Argon vs GPT-6 Astra Pricing
Your choice becomes more interesting when comparing Argon with OpenAI’s GPT-6 Astra.
Elite Era Trends’ current analysis of Astra lists standard pricing of:
$10 per million input tokens
and:
$50 per million output tokens
for GPT-6 Astra.
You can see the full breakdown in our GPT-6 Astra pricing, features and context-window guide.
Using those published rates, Gemini 4 Argon’s introductory token prices are substantially lower.
Simplified comparison
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Gemini 4 Argon — intro | $2 | $10 |
| Gemini 4 Argon — Jan. 2027 | $4 | $20 |
| GPT-6 Astra standard | $10 | $50 |
This comparison considers only standard token pricing.
It does not prove Argon is the better or cheaper model for every actual task.
Overall cost can depend on:
- number of calls;
- token usage;
- caching;
- tool use;
- reasoning behavior;
- output length;
- retries; and
- whether the model successfully completes the task.
For information about how much Astra usage is available through ChatGPT plans, see our GPT-6 Astra usage limits guide.
Gemini 4 Argon vs GPT-6 Astra Cost Example
Suppose an API workload uses:
1 million input tokens
and:
100,000 output tokens
Gemini 4 Argon introductory rate
Input:
$2
Output:
$1
Total:
$3
Gemini 4 Argon 2027 rate
Input:
$4
Output:
$2
Total:
$6
GPT-6 Astra at Elite Era Trends’ currently documented standard rate
Input:
$10
Output:
100,000 ÷ 1M × $50 = $5
Total:
$15
For this specific token-count example:
Argon intro: $3
Argon 2027: $6
Astra: $15
But price alone is not enough to determine which model produces the best economic result.
A model requiring three attempts to complete a task can cost more than a more expensive model that succeeds on the first try.
Output Tokens Can Dominate Your Bill
One of the most important details in Argon’s price structure is the difference between input and output.
Output is priced at:
5× the input-token rate
during both announced pricing periods.
During launch pricing:
$2 input vs $10 output
From January 2027:
$4 input vs $20 output
Therefore, controlling unnecessary output can materially reduce costs.
For example, at the introductory rate:
1 million input + 10,000 output
costs:
$2 + $0.10 = $2.10
But:
1 million input + 1 million output
costs:
$2 + $10 = $12
The input is identical.
The large difference comes from output generation.
How to Reduce Gemini 4 Argon API Costs
1. Use Argon only for difficult tasks
Reserve frontier-model calls for work where advanced reasoning creates real value.
Simple extraction, classification or formatting may not require Argon.
2. Limit unnecessary output
Request the shortest output that still satisfies the task.
If you need JSON with five fields, do not ask the model to produce a five-page explanation first.
3. Use caching when appropriate
Repeated large context can become expensive.
Google’s announced 95% cached-input discount could materially lower repeated-context costs.
4. Break workflows into model tiers
A practical architecture could be:
Cheap model → filter/classify → Argon only when needed
rather than:
Argon handles every request
5. Monitor token usage
Log:
- input tokens;
- output tokens;
- cached tokens;
- calls per user;
- cost per workflow; and
- successful task completion rate.
The important metric is often:
Cost per completed task
rather than merely:
Cost per token
6. Design prompts efficiently
Sending unnecessary documentation with every request increases input usage.
Use retrieval or selective context when possible.
7. Budget using 2027 rates
If a product will operate beyond December 2026, build financial forecasts around:
$4 input / $20 output
rather than assuming introductory prices remain permanent.
What Could 1,000 Users Cost?
Imagine a product has:
1,000 users
Each user generates:
100,000 input tokens
and:
20,000 output tokens
per day.
Cost per user at introductory pricing:
$0.40/day
For 1,000 users:
$400/day
Across 30 days:
$12,000/month
At the announced 2027 rate:
$0.80 per user/day
For 1,000 users:
$800/day
Across 30 days:
$24,000/month
Again, this is a simplified token-only calculation.
But it demonstrates why token economics become important once an AI application has meaningful scale.
What Would 1 Billion Argon Tokens Cost?
This depends on the mix of input and output.
1 billion input tokens at introductory price
1 billion tokens = 1,000 million-token units.
1,000 × $2 = $2,000
1 billion output tokens
1,000 × $10 = $10,000
From January 2027
1 billion input tokens:
1,000 × $4 = $4,000
1 billion output tokens:
1,000 × $20 = $20,000
This illustrates why output-heavy products can have very different economics from retrieval-heavy or analysis-heavy products.
Does Google AI Ultra Include Free Gemini 4 Argon API Tokens?
Google has announced Google AI Ultra subscribers as an early consumer-access group for Argon.
However, Google’s Argon launch announcement does not state that an AI Ultra subscription includes a particular quantity of free Gemini 4 Argon API tokens.
Consumer subscription access and developer API billing should therefore be treated as separate products unless Google explicitly documents otherwise.
If you are trying to determine who can access Argon now, see our Gemini 4 Argon availability guide.
When Does Gemini 4 Argon’s Introductory Pricing End?
Google’s official regional launch materials state that the introductory pricing applies through:
December 31, 2026
Beginning:
January 1, 2027
the announced rates become:
$4 per million input tokens
and:
$20 per million output tokens
Developers launching products late in 2026 should therefore model their long-term economics using the higher rate.
A workload that appears profitable at $2/$10 could look quite different at $4/$20.
Is Gemini 4 Argon Available Through the API Yet?
Google says paid API customers will be among the first broader groups to receive Gemini 4 Argon.
However, as of October 6, the model remains in a phased rollout.
Google has not announced a universal API-availability date.
That means developers should verify that Argon is actually available in their account before building a production dependency around it.
Check Google’s official model documentation rather than relying on an unofficial model name or social-media screenshot.
Frequently Asked Questions
How much does Gemini 4 Argon cost?
Google has announced introductory pricing of $2 per million input tokens and $10 per million output tokens.
When does the Gemini 4 Argon introductory price end?
Google’s official launch materials state that the introductory period runs through December 31, 2026.
What will Gemini 4 Argon cost in 2027?
Beginning January 1, 2027, Google says the price will be $4 per million input tokens and $20 per million output tokens.
How much do 1 million Gemini 4 Argon tokens cost?
It depends on whether they are input or output tokens.
During introductory pricing:
1M input = $2
1M output = $10
From January 2027:
1M input = $4
1M output = $20
How much do cached Gemini 4 Argon tokens cost?
Google says cached input receives a 95% discount from the input-token price. During the $2 introductory input rate, 95% off corresponds to approximately $0.10 per million cached input tokens, assuming the tokens qualify under Google’s caching rules.
Is Gemini 4 Argon cheaper than GPT-6 Astra?
Using the currently published standard token rates, Argon’s announced rates are lower than GPT-6 Astra’s documented standard token price. Actual workload cost depends on token use, model behavior, tools, retries and task success.
Does Gemini 4 Argon have a free API tier?
Google has not announced a general free Argon API tier in its launch post.
Can Google AI Ultra users use the Argon API for free?
Google has not stated that AI Ultra consumer subscriptions include a specific allowance of free Argon API tokens.
Is Gemini 4 Argon available now?
Argon remains in a phased rollout. Google says broader access will begin with paid API customers and Google AI Ultra subscribers, but no universal release date has been announced.
Why are output tokens more expensive?
Generating model output consumes inference resources. Google prices Argon output at five times the corresponding input-token rate under both announced pricing periods.
Bottom Line
Gemini 4 Argon launches with aggressive frontier-model API pricing.
Through December 31, 2026, Google says Argon will cost:
$2 per million input tokens
and:
$10 per million output tokens
From January 1, 2027, those prices are scheduled to double to:
$4 per million input tokens
and:
$20 per million output tokens
Google also announced a 95% discount on cached input tokens, making caching potentially important for large repeated-context workloads.
For developers, the main lesson is to avoid evaluating Argon only at its introductory rate.
If an application is intended to operate into 2027, its financial model should be tested against the $4/$20 rates.
And because output is priced at five times input, controlling unnecessary response length can make a meaningful difference at scale.
Argon’s token price is only one part of the decision.
The real metric is:
How much does the model cost to successfully complete the job?
A cheaper model that requires more retries or produces lower-quality results can ultimately cost more than a higher-priced model that reliably completes difficult work.
Related Reading on Elite Era Trends
How to Access Gemini 4 Argon: Availability, AI Ultra, API & Fairwind Explained
GPT-6 Astra: Pricing, Features, Context Window & Complete 2026 Guide
GPT-6 Astra Usage Limits Explained
GPT-6 Astra for Work: Features, Computer Use & Best Use Cases