API budgeting guide
How to Estimate API Cost Without Confusing It with a Subscription
By Image to Video Prompt · Published
An API cost calculator answers two questions: what does one request cost, and what will your whole workload cost? Start with tokens per request, then multiply by the number of requests. Keep that usage budget separate from a chat subscription. This guide walks through a reproducible example before you enter your workload in our AI Model API Cost Calculator.
1. Measure one request before scaling up
Input tokens are the billable material sent to the model; output tokens are the billable material it produces. Use the provider’s usage counts when available. Your visible prompt may omit system instructions, conversation history and tool schemas, while billable output may include thinking tokens. A word count is not an exact token count.
In the calculator, “Input tokens per call” and “Output tokens per call” describe one request. Check the unit selector: 2,000 raw tokens, 2K and 0.002M mean the same input. Entering 2,000 with the M unit would mean two billion tokens. The initial 1M input and 1M output are editable examples, not a typical request size or a provider allowance.
Count API requests rather than users or finished tasks. One task might call a model repeatedly. If request sizes differ substantially, estimate each group separately and add the totals. Repeating an average request is useful only when it represents the workload.
2. Calculate the single-request cost
For this example only, assume $3 per million input tokens and $12 per million output tokens, with no discounts or request tiers. These are hypothetical teaching rates, not a current vendor quote or a selectable model in our calculator. One request uses 2,000 input tokens and 500 output tokens.
Input: 2,000 ÷ 1,000,000 × $3 = $0.006
Output: 500 ÷ 1,000,000 × $12 = $0.006
Per request: $0.006 + $0.006 = $0.012
Keep input and output calculations separate because their rates can differ. Charging all 2,500 tokens at the input rate would underestimate this example. Keep the small decimals until the final total; rounding each request to the nearest cent can distort a large workload.
3. Multiply by calls, then compare with your budget
At the same hypothetical rates and token counts, 100 requests cost $1.20 and 1,000 cost $12. Changing the call count changes the total, while the single-request cost stays $0.012.
| Calls | Per call | Total |
|---|---|---|
| 1 | $0.012 | $0.012 |
| 100 | $0.012 | $1.20 |
| 1,000 | $0.012 | $12.00 |
A $30 token budget covers 2,500 identical requests in this simplified example: $30 ÷ $0.012. Reserving $5 for uncertainty leaves $25 for 2,083 whole requests. That is a planning estimate, not an enforced spending cap. Our calculator does not stop requests or manage a provider account.
Request tiers are separate. Where the selected model charges more above an input-length threshold, the tool checks tokens in one request. Repeating a small request many times does not make it one long request. Increasing tokens per call can change the tier; increasing the call count alone does not.
4. Keep subscriptions and API usage in separate columns
A monthly chat plan price is not a per-token API rate or a fixed number of API requests. Anthropic’s subscription and API billing explanation describes them as separate products. As checked October 12, 2026, eligible Max and Team plans also offer monthly API credits with claim and expiry conditions. Pro is not eligible for that credit. Confirm any credit in your linked Console account before treating it as available.
Estimate gross API usage first, then account for confirmed credits separately. Our tool does not read your balance or apply subscription credits. Its Pro/Max comparison shows monthly plan prices; it does not convert them into the hypothetical request allowance above. If you use both products, record the subscription fee alongside your API budget.
5. Use the right estimate for the job
Open the calculator’s token budget, select an included model, and check its linked official source and verification date. Enter input and output per call with the correct units, then your call count. Read both per-call and total results. Pricing can change after the displayed check date.
The separate image budget uses image count and resolution, with Standard or asynchronous Batch options. Video duration and reference inputs belong in the Seedance prompt and video budget planner. Neither uses the hypothetical text-token rates in this guide.
Include only discounts explicitly modeled by your selection. DeepSeek has manual peak/off-peak and input-cache assumptions; other token entries do not share those settings. Tools, taxes, reseller fees and billable retries can add costs. Follow the selected provider’s official source; for example, check Claude pricing and feature charges and reconcile actual usage afterward. The provider’s usage record and invoice establish what you owe.
