Free browser tool
LLM Cost Calculator
LLM cost calculator for API input, output, cached tokens and batch assumptions using rates you enter and verify.
LLM cost calculator: compare user-entered token rates and billing assumptions for a measured workload.
API tokens, cache and batch factors
Enter total input tokens, output tokens, cached input tokens and separate input, output and cached-input rates per million. Cached tokens cannot exceed total input. Batch-rate ratio multiplies the token charge only: 1 means no discount, and 0.5 is an illustrative half-rate assumption. Verify whether your selected provider supports that assumption.
All rates and costs must use the same currency. Example amounts are fictional assumptions. No vendor rates are preselected and no API calls are made.
Worked example
For 1,000,000 input tokens including 250,000 cached tokens, 200,000 output tokens, illustrative rates of 1 input, 0.25 cached input and 4 output, input cost is 0.8125 and output cost is 0.8. Total cost before any batch multiplier is 1.6125.
Assumptions and limits
Input tokens must include the cached subset only once. Provider invoices can include cache-write charges, reasoning tokens, tools, minimums or other items that this simple token scenario does not model. Add those to your TCO record and validate rates against the official page.
Pricing references checked 6 October 2026
Use the linked official provider pricing pages for the model and service you actually select. Rates vary by model, processing mode, cache operations and service. Copy the applicable rates into your assumptions and retain the source and date. This calculator’s illustrative rates are not a model price table.
Who should use this calculator and how to read it
Use this page when you want to isolate API token charges for a particular service or model. Enter the total input and output token counts, the cached portion of input, and the corresponding rates per million in one currency. Cached tokens are part of total input; the calculation separates them so the same tokens are not billed twice. A batch-rate ratio applies only when your provider and request type qualify.
Keep the selected model, processing mode, region, rate source, and date beside your result. Check the service's actual invoice rules for cache writes, reasoning categories, tools, minimums, and other fees. The agent cost calculator adds tasks, steps, retries, tools, and operating costs; the token calculator focuses on measured counts; the ROI calculator models time value separately.
Frequently asked questions
Can cached tokens exceed total input tokens?
No. The cached count is a subset of input and the calculator rejects a cached count larger than the input count. Include each input token once, under either its cached or uncached rate.
Does the batch factor guarantee a discount?
No. It is a multiplier you enter for a scenario. Check the current provider terms, eligible endpoint, model, and workload before using a reduced factor. A value of one applies no reduction.
Does the estimate include every invoice item?
No. It models the token charges and batch factor shown on this page. Add any applicable cache-write, tool, hosting, minimum, tax, or service fees separately using documented billing units.
Does this page retrieve rates or send prompts to a model?
No. It accepts numeric assumptions in your browser and makes no model request or price lookup. Use measured token counts and the current official price source or contract for your actual service.
Compare equivalent API workloads
An LLM API cost calculator or LLM pricing calculator should compare the same model workload, currency, region, processing mode and billing period. An OpenAI API cost calculator, Claude API cost calculator, Claude API pricing calculator, Gemini API cost calculator or GPT API pricing calculator is useful only when you enter the applicable current terms. The official Gemini API pricing page is one source to check; the linked official OpenAI and Anthropic pages are also provided above. Prices and eligible features can change.
For an LLM cost comparison or LLM pricing comparison, separate input, output and cached-input units before totaling. In the fictional sample, 750,000 uncached input tokens at $1 per million cost $0.75; 250,000 cached tokens at $0.25 per million cost $0.0625; and 200,000 output tokens at $4 per million cost $0.80. Total is $1.6125 before batch adjustment. A cost per million tokens comparison should keep input and output rates distinct.
A prompt caching savings calculator scenario can estimate $0.1875 saved on 250,000 tokens at fictional rates of $1 uncached and $0.25 cached: 250,000 × ($1 − $0.25) ÷ 1,000,000. A batch API cost savings calculator scenario may test an entered batch-rate ratio, but a value such as 0.5 is only an assumption. Confirm availability, eligible models and terms in the provider’s current documentation before treating it as a discount. This page does not identify the cheapest LLM API. To compare LLM API prices, use the same workload and review current terms rather than relying on one unit rate.
Updated 2026-10-08. Sources are linked on this page.
Primary sources and review
- OpenAI: official API pricing and cost units
- Anthropic: official pricing, caching and batch rates
- NIST AI RMF: risk-management work to account for in operating costs
Published by AI Agent Cost Calculator. Last updated: . Methods on this site are practical workflows; outputs do not certify compliance.