LLM Memory Calculator

Size the hardware you need to serve open-weight LLMs. Estimate VRAM for weights, KV cache growth per token and total memory at your target concurrency.

Last updated . Model pricing is refreshed twice daily.

What it does

Size the hardware you need to serve open-weight LLMs. Estimate VRAM for weights, KV cache growth per token and total memory at your target concurrency.

  • Weight, activation and KV cache breakdown per model
  • Quantization support (FP16, INT8, INT4, GPTQ, AWQ)
  • Concurrency and batch-size aware
  • Recommends GPU class and count

Why it matters

Tokenomy is the economic runtime for AI — it shows where every AI dollar goes and gives you the controls to redirect it. This tool is free to use; sign in to save scenarios, share with your team and connect it to your live usage ledger.