Writing

Working Out What an Agent Will Cost

The estimate changes with workload, review time, and where each task runs

Two lanes Some workloads can run locally; others need a different model or a different review process.

A simple enterprise AI business case might begin like this. Pick a model, estimate a volume, multiply by a published price, compare against the labor it replaces. The number comes out favorable, the program gets funded, and the actual bill arrives shaped differently.

That estimate leaves out the decisions about routing work between models, along with cost variables that can change independently. Research has examined several approaches: cascading cheap models ahead of expensive ones changes the cost curve substantially (Chen, Zaharia & Zou, 2023), and per-query routing by expected difficulty does the same (Ding et al., 2024).

I would want the estimate to show where different workloads will run, who decides, and how much work a later change would require.

Variables to include

Price per token is easy to compare. These other variables also belong in the estimate, particularly as more people begin using the system.

Volume grows with adoption, non-linearly. A pilot serving a team is rarely a tenth of the cost of one serving the department, because the department finds uses the pilot team never proposed. Include the possibility that a useful pilot attracts additional work.

Agentic work multiplies calls per task. A single user request is rarely a single call. This is the ongoing-cost half of what the machine learning debt literature warned about: the system around the model is where the recurring expense lives (Sculley et al., 2015). It is retrieval, then reasoning, then a tool call, then a check, sometimes a retry after a failure. The multiplier between "requests" and "calls" is set by the architecture, is easy to change accidentally, and is frequently absent from the model entirely. A modest change in retry policy can move the bill more than switching providers.

Context size affects every call. Systems tend to send far more context than the task needs, because sending everything is easier than deciding what matters. Reducing unnecessary context can lower recurring costs, provided the task still receives the information it needs.

The variables that actually move the number Track the calls, context, review time, and capacity each workload uses.

Human review time is a real cost and belongs in the same model. Programs that book the inference saving and ignore the organizational cost repeat one of the best documented errors in the productivity literature (Brynjolfsson & Hitt, 2000). If an output replaces thirty minutes of work but needs six minutes of review, the time saving is twenty-four minutes. The financial saving also depends on whose time is involved.

Running some work locally

Running a model on your own hardware can be cheaper for some workloads. The comparison depends on current prices, capacity, and the quality required. Capability per unit of compute keeps moving on its own curve (Stanford Institute for Human-Centered AI, 2026).

Owning hardware shifts part of the cost into capacity purchased in advance. Additional calls still consume electricity, time, and available capacity, but there is no provider invoice for each one. Under per-call pricing, running a check twice, or running three approaches and comparing, or reprocessing the whole back catalogue because you improved a prompt, are all things you weigh. With spare local capacity, those experiments may be easier to justify.

That shift matters most for exactly the high volume, low judgment work that dominates real enterprise usage: classification, extraction, routing, summarizing, first-pass reconciliation. The working systems I see are dominated by this category rather than by headline reasoning. Work where the output is checked anyway, where the difficulty is low and the volume is enormous.

There is a second property that is not about money at all. Data that never leaves your control is a different conversation with legal, procurement, and any customer contract with a data residency clause. For some workloads, those requirements may decide where the data can be processed before price enters the discussion.

Owning that capacity also means paying for maintenance and eventual replacement. Include the people needed to operate it when comparing the cost with a hosted service.

The routing question

Given both lanes, the design question is which work goes where, and I would consider tolerance for a wrong answer and volume alongside measured task performance.

High volume, high tolerance, low judgment work is the natural local lane. Wrong answers are caught downstream, the value is in the aggregate, and the economics reward doing far more of it than a per-call budget would permit.

Low volume, low tolerance, high judgment work is the natural frontier lane. The call count is small, the potential cost of an error may outweigh the inference bill, and you want the best available reasoning on the problem. Deciding which model is actually better for a given task means measuring across scenarios rather than trusting a single headline number (Liang et al., 2022).

The dangerous quadrant is high volume, low tolerance, which is where teams reach for the frontier model and then discover the bill. That combination is usually a signal that the process needs redesigning rather than routing: the required quality may need a different workflow, additional checks, or a narrower scope.

What to build regardless of the answer

Build the routing seam on day one, even if it points at a single provider today. An abstraction over model selection costs little at the start and is very expensive to retrofit into a codebase that assumed one vendor's response shape.

Measure per-workload, not per-month. A single invoice can hide which work is driving the cost. Treat the instrumentation as part of production readiness rather than as reporting (Breck et al., 2017). Cost attributed to a workload tells you which use case is quietly consuming the budget, and in my experience it is almost never the one people assume.

Instrument the multiplier. Track calls per completed task. It is the number most likely to drift upward without anybody deciding, and the one most responsive to a day of attention.

Document the fallback too. If a provider is unavailable, rate limits you, deprecates the model you built on, or changes its pricing mid-year, what happens? Answering that shapes the architecture more than any current price does, because all four of those are ordinary events rather than tail risks.

Every crossover number expires

A crossover estimate, the volume at which local becomes cheaper than a frontier service, needs a date and its assumptions attached. The original scaling relationships (Kaplan et al., 2020) were themselves revised once the compute-optimal balance was reexamined (Hoffmann et al., 2022), and that revision moved the economics of serving. Model prices have moved repeatedly, open weight capability has moved faster, and hardware moves on its own schedule. Revisit the estimate when those inputs change.

Separate fixed and recurring costs, track calls per task, and record the quality required for each workload. The mapping and management functions in governance frameworks provide a useful place to maintain those decisions (National Institute of Standards and Technology, 2023). I would keep the assumptions beside the calculation and assign someone to revisit them as usage grows.

The switch Keep the routing decision easy to revise as costs and requirements change.

References

  1. Chen, Zaharia & Zou (2023). FrugalGPT: How to use large language models while reducing cost and improving performance. arXiv:2305.05176. arxiv.org/abs/2305.05176 Cascading cheaper models before expensive ones, with the cost and quality tradeoff measured.
  2. Ding et al. (2024). Hybrid LLM: Cost-efficient and quality-aware query routing. arXiv:2404.14618. arxiv.org/abs/2404.14618 Query routing between a small local model and a large one, decided per query by expected difficulty.
  3. Kaplan et al. (2020). Scaling laws for neural language models. arXiv:2001.08361. arxiv.org/abs/2001.08361 Performance as a smooth function of compute, data and parameters.
  4. Hoffmann et al. (2022). Training compute-optimal large language models. arXiv:2203.15556. arxiv.org/abs/2203.15556 Revised the compute-optimal balance, and with it the economics of serving.
  5. Stanford Institute for Human-Centered AI (2026). AI Index Report. Stanford University. hai.stanford.edu/ai-index
  6. Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28. proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
  7. Liang et al. (2022). Holistic Evaluation of Language Models. arXiv:2211.09110. arxiv.org/abs/2211.09110
  8. Brynjolfsson & Hitt (2000). Beyond computation: Information technology, organizational transformation and business performance. Journal of Economic Perspectives, 14(4), 23-48. doi.org/10.1257/jep.14.4.23 Why measured returns to information technology depend on complementary organizational change.
  9. National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
  10. Breck et al. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. IEEE International Conference on Big Data. doi.org/10.1109/BigData.2017.8258038