NEWTON’S RESEARCH · INFERENCE ECONOMICS

The Same AI Workload Doesn’t Have One Price

AI inference is not sold at one universal price. The economically relevant unit is a workload with capability, latency, reliability, context, and routing constraints — not a generic token detached from how it is served.

By Newton’s ResearchPublished August 21, 2026Last verified August 21, 20269 min readResearch methodology

What the evidence establishes

Evidence is separated from inference and hypothesis so the conclusion can be challenged without changing the facts.

Major model providers publish multiple prices or service modes for input, cached input, output, batch or lower-priority processing, and other workload characteristics.

Gateway products already expose provider selection, budgets, fallbacks, and routing controls because price is only one dimension of the decision.

The same nominal model family can be operationally different across routes because latency, feature support, data handling, availability, and caching behavior can differ.

Evidence — observed or documented facts.Inference — what follows from those facts.Hypothesis — what remains to be tested.

A price-per-million-tokens table looks like a commodity quote, but it leaves out the variables that determine whether the workload actually finishes well. A builder is buying an outcome under constraints: a model must be capable enough, available soon enough, compatible with the toolchain, and predictable enough to trust with the task.

Evidence: providers already publish a price function

OpenAI, Anthropic, and Google distinguish among billing categories such as ordinary input, cached context, output, and alternative processing modes. Those categories exist because serving the same broad class of work under different constraints has different economics. This is visible in first-party pricing before a gateway enters the picture.

Inference: abstraction becomes more valuable as the matrix grows

Once a developer uses several model families and several economic modes, the decision stops being a simple model-choice problem. It becomes a portfolio problem: which route satisfies the task at the lowest acceptable generalized cost? Existing gateways already expose price, provider, fallback, budget, and routing controls because this coordination problem is real.

  • A low token price is irrelevant if the model needs expensive retries.
  • A fast route may be worth more when a deadline makes delay expensive.
  • A cached workflow may have different economics from a fresh long-context workflow.
  • A direct provider can be rational when one model dominates and intermediary risk matters more than portability.

Hypothesis: inference may become increasingly market-like

The hypothesis is not that every token becomes interchangeable. Models produce different work and providers can expose different operational characteristics. The narrower claim is that more of the access layer may behave like a market: buyers compare multiple sellers, service qualities, commitments, and execution paths for machine intelligence.

Do not accept the conclusion on faith.

Use the decision rule against your own workload. Newton’s should win the next request only when the evidence says it should.