Your privacy choices

Allow optional cookies for referral attribution, visit analytics, and Google Ads purchase measurement.

Pricing calculator

Estimate your LLM API cost before you commit

Rates come from the current public catalog. Input and output tokens are billed separately.

Input / 1M¥1.20
Output / 1M¥6.00
Cached input / 1M¥0.12
Effective rate per 1M¥3.60
Estimated monthly cost¥360.00

CNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.

How this calculator differs from guessing

Rates come from the current public catalog. Input and output tokens are billed separately.

This calculatorA rough estimate
ModelRates come from the current public catalog. Input and output tokens are billed separately.Memory or a blog post
Cache-hit rateCache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount.Usually ignored

Who this is for

Useful if you

  • Want a defensible cost estimate before moving traffic to a model
  • Want to see how a higher cache-hit rate lowers your effective rate
  • Are budgeting a monthly LLM spend in CNY

Less useful if you

  • Need USD figures — this calculator works in CNY
  • Bill per image or per call rather than per token (those rows are excluded)
  • Want a guaranteed quote — actual usage depends on your prompts and outputs
  • Need provider-native pricing for a model the gateway does not route

How to read the result

The math is simple and transparent, but estimates are estimates.

  • Cost = uncached input × input price + cached input × cache price + output × output price.
  • Cache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount.
  • Only token-priced models with published input and output prices are included; per-call models are excluded.
  • CNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.

The rule behind the numbers

Rates come from the current public catalog. Input and output tokens are billed separately.

Input / 1MCNY / 1MCost = uncached input × input price + cached input × cache price + output × output price.
Output / 1MCNY / 1MCNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.
Cached input / 1MCNY / 1MCache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount.

This is an estimate based on published rates and your inputs, not a quote. Actual cost depends on real token usage and cache behavior in your workload.

Pricing calculator — FAQ

Where do the rates come from?

Rates come from the current public catalog. Input and output tokens are billed separately.

What does the cache-hit rate do?

Cache-hit percentage applies only to input, using the published cache price. No cache price means no assumed discount.

Is the result a guaranteed price?

No. It is an estimate. Your real cost depends on actual input and output token counts and how much of your input is cached.

Why is everything in CNY?

CNY per million tokens. This estimates metered usage, not the cash price of a prepaid package.

Happy with the estimate?

Buy a key and start — prepaid credits, no expiry.