Lesson 11 / 25

Running Costs

Tokens, services, infrastructure.

Measure tokens per item

Running costs scale with volume: model calls (input and output tokens times price, times calls per item), other services such as OCR or document parsing, and infrastructure (hosting, queues, storage, logging, monitoring). Measure tokens per item in the pilot, include retries and long documents, and check provider prices and volume discounts. For many back-office automations, running costs are modest compared with labour; confirm it with your numbers.

Monthly running cost, run

I ran this with plain Python 3. All figures belong to one worked example (an invoice-processing team) with invented but internally consistent numbers; prices are placeholders. Two model calls per invoice with 3,500 input and 400 output tokens each cost about $396 a month at placeholder prices; OCR adds $120 and infrastructure $400, for $916 a month or about $0.076 per invoice.

volume = 12000
calls_per_item = 2                  # extraction + validation
tokens_in, tokens_out = 3500, 400   # per call, measured on samples
price_in, price_out = 3.00, 15.00   # $ per million tokens (placeholder)
llm = volume * calls_per_item * (tokens_in * price_in + tokens_out * price_out) / 1e6
ocr = volume * 0.01                 # document parsing service, $ per page (placeholder)
infra = 400                         # hosting, queue, logs, monitoring per month
print(f"LLM calls  ${llm:>8,.0f}/month")
print(f"OCR        ${ocr:>8,.0f}/month")
print(f"infra      ${infra:>8,.0f}/month")
print(f"total run  ${llm + ocr + infra:>8,.0f}/month  (${(llm + ocr + infra) / volume:.3f} per invoice)")

Output:

LLM calls  $     396/month
OCR        $     120/month
infra      $     400/month
total run  $     916/month  ($0.076 per invoice)

Use pilot token counts

Real documents are longer and messier than samples; measure tokens on production-like data.

Quick check: How should model costs be estimated?

  • They are always zero
  • From a guess of one cent per month
  • From the number of engineers
  • From measured tokens per item, calls per item and volume
Answer

From measured tokens per item, calls per item and volume — Measure and multiply.