Estimate your LLM API cost before you commit
Rates come from the current public catalog. Input and output tokens are billed separately.
CNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.
How this calculator differs from guessing
Rates come from the current public catalog. Input and output tokens are billed separately.
| This calculator | A rough estimate | |
|---|---|---|
| Model | Rates come from the current public catalog. Input and output tokens are billed separately. | Memory or a blog post |
| Cache-hit rate | Cache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount. | Usually ignored |
Who this is for
Useful if you
- Want a defensible cost estimate before moving traffic to a model
- Want to see how a higher cache-hit rate lowers your effective rate
- Are budgeting a monthly LLM spend in CNY
Less useful if you
- Need USD figures — this calculator works in CNY
- Bill per image or per call rather than per token (those rows are excluded)
- Want a guaranteed quote — actual usage depends on your prompts and outputs
- Need provider-native pricing for a model the gateway does not route
How to read the result
The math is simple and transparent, but estimates are estimates.
- Cost = uncached input × input price + cached input × cache price + output × output price.
- Cache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount.
- Only token-priced models with published input and output prices are included; per-call models are excluded.
- CNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.
The rule behind the numbers
Rates come from the current public catalog. Input and output tokens are billed separately.
This is an estimate based on published rates and your inputs, not a quote. Actual cost depends on real token usage and cache behavior in your workload.
Pricing calculator — FAQ
Where do the rates come from?
Rates come from the current public catalog. Input and output tokens are billed separately.
What does the cache-hit rate do?
Cache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount.
Is the result a guaranteed price?
No. It is an estimate. Your real cost depends on actual input and output token counts and how much of your input is cached.
Why is everything in CNY?
CNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.
Happy with the estimate?
Buy a key and start — prepaid credits, no expiry.