This guide groups direct usage and operating effort into a repeatable estimate. Enter the rates for the provider, model, service region, and contract you actually use; provider pricing changes and billing units differ. The calculator linked below accepts your assumptions and runs in your browser. It does not fetch live prices or reproduce every provider invoice line.
Model cost
Start with observed tasks in a month, average model steps per task, and input and output tokens per step. Split cached tokens or other billable token classes only when the provider reports them and the calculator mode supports them. Include retries as additional work, not as a vague contingency percentage. Record model, endpoint, currency, rate unit, source URL, and retrieval date beside each rate. The official OpenAI API pricing and Google Agent Platform pricing pages illustrate why the service, model, processing mode, and extra features must be checked directly. Rates on those pages are not interchangeable.
Illustrative example: suppose a workflow runs 1,000 tasks/month, averages 4 model steps/task, 1,500 input and 300 output tokens/step, and uses user-entered rates of $1 and $4 per million tokens. Token charges are (1,000 × 4 × 1,500 ÷ 1,000,000 × $1) + (1,000 × 4 × 300 ÷ 1,000,000 × $4) = $8.80/month. These are made-up arithmetic inputs, not a provider quote. Tool fees and hosting are excluded from this subtotal.
Orchestration and tools
Count charges outside model tokens: hosted execution time, search or retrieval calls, storage, network transfer, third-party APIs, and orchestration infrastructure. Keep fixed monthly commitments separate from usage-based charges. If a provider prices a tool call, container session, or data store separately, model that line with its own quantity and unit rate. Do not count a tool in both the model rate and a separate tool subtotal.
For the example above, suppose the workflow makes three external tool calls per task at an assumed $0.002 per call, plus $25/month of allocated hosting. Tool usage is 1,000 × 3 × $0.002 = $6/month, bringing the known subtotal to $39.80 ($8.80 model + $6 tools + $25 hosting). The $25 and $0.002 are fictional inputs.
Observability
Include traces, logs, evaluation runs, monitoring, and retained test data when they create incremental cost. Separate production traffic from offline evaluation so an increase can be explained. Track both cost per completed task and the underlying drivers—tokens, tool calls, retries, and evaluator runs. The FinOps Foundation’s Unit Economics capability describes relating technology spend to an appropriate business unit and reviewing the metric over time. Choose a denominator that represents useful work, such as completed task or resolved case, and define what qualifies as complete.
Security and governance
Budget the people and services actually required for access reviews, data handling, evaluation, change approval, incident response, and audit evidence. Do not assume that a particular framework mandates a fixed percentage or specific cost line: no authorized normative text was used to derive such a rule here. Record which costs are direct invoices and which are internal allocations. For sensitive workflows, include the cost of required human review and safe failure handling rather than treating them as optional overhead.
People
Separate initial design and integration effort from recurring operations. For recurring effort, use hours × your organization’s loaded hourly rate, and state whether that number is cash expenditure, allocated staff capacity, or an opportunity-cost estimate. Avoid counting the same engineer hours under both project labor and support. The FinOps Foundation defines TCO broadly to include management, support, labor, and other costs; its terminology is a useful prompt for inclusions, not a universal allocation formula.
Continue the example with an assumed 10 support/review hours at $50/hour ($500) and $40/month for monitoring. The illustrative monthly total becomes $579.80. It is an estimate of the stated scope, not a forecast or a promised saving. If the internal labor is not incremental cash, keep it visible as allocated effort rather than presenting the total as a new bill.
Template
Use this compact record for each estimate and keep the detailed evidence with it:
| quantity × rate | monthly amount | source / owner | |
|---|---|---|---|
| Model input and output | tasks × steps × tokens × entered rates | your calculation | provider price page + retrieval date |
| Tools and orchestration | calls, runtime, storage, fixed fees | your calculation | service billing page or invoice |
| Observability and evaluation | events, runs, retention | your calculation | invoice / internal owner |
| People and review | hours × loaded rate | your calculation | team estimate; label allocation |
| One-time implementation | hours and external fees | keep separate from monthly TCO | estimate owner and date |
Reconcile the estimate against an actual billing period before using it for a budget decision. Compare like periods and the same scope; explain changes in volume, rates, retries, and included work. The AI agent cost calculator covers a user-entered model, tool, fixed-cost, and human-cost scenario; use the LLM API cost calculator for token and cached-input assumptions and the token cost calculator for measured counts. None pulls live rates or sends your inputs to an API.