The same model costs different money on different hosts

The same open-weight model is served by multiple hosts at prices that commonly differ by two to five times for identical weights. Tokenomy tracks every host serving each model, with input and output price per million tokens, recent uptime and context window, so you can see the cheapest qualified host before you commit traffic.

Open-weight models are served by many clouds at once. The weights are identical; the bill is not. Every multi-host model we track, ranked by how far apart its hosts price it.

Last updated . Model pricing is refreshed twice daily.

Why one price per model is misleading

Most price tables list a single rate per model. For open-weight models that number is an average of very different offers: quantization, context limits, throughput and margin all vary by host.

  • Identical weights, different hosts, multiples of price difference
  • Blended at three input tokens per output token, the mix most production traffic runs
  • Recent uptime shown next to price, because cheapest is useless if it is down
  • Quantization noted, because cheaper sometimes means a smaller-precision build

From comparison to control

Knowing the cheapest host is step one. Tokenomy's router moves traffic to the cheapest option that still clears your quality and latency bar, and the ledger proves what the move saved.

Frequently asked questions

Why does the same model cost different amounts on different providers?

Hosts buy different hardware, run different quantizations and take different margins. For open-weight models nothing stops several clouds serving the same weights at their own price.

Is the cheapest host always the right choice?

No. Check quantization, context limit and recent uptime. A lower-precision build at a lower price can change output quality, so validate on your own evaluation set.

How often are the endpoint prices refreshed?

Daily, from the hosts' published endpoint data, with the observation time shown on the page.