LLM Memory Calculator
Size the hardware you need to serve open-weight LLMs. Estimate VRAM for weights, KV cache growth per token and total memory at your target concurrency.
What it does
Size the hardware you need to serve open-weight LLMs. Estimate VRAM for weights, KV cache growth per token and total memory at your target concurrency.
- Weight, activation and KV cache breakdown per model
- Quantization support (FP16, INT8, INT4, GPTQ, AWQ)
- Concurrency and batch-size aware
- Recommends GPU class and count
Why it matters
Tokenomy is FinOps for AI — see where every AI dollar goes and get 20–40% of them back. This tool is free to use; sign in to save scenarios, share with your team and connect it to your live usage ledger.