# LLM Memory Calculator

> Calculate VRAM, KV cache and system memory for self-hosted LLM inference. Size hardware for Llama, Mistral, Qwen, DeepSeek and other open-weight models.

Source: https://tokenomy.ai/tools/memory-calculator
Last updated: 2026-09-22
Publisher: Tokenomy — FinOps for AI
License: free to quote with attribution and a link to the source URL.

Size the hardware you need to serve open-weight LLMs. Estimate VRAM for weights, KV cache growth per token and total memory at your target concurrency.

## What it does

Size the hardware you need to serve open-weight LLMs. Estimate VRAM for weights, KV cache growth per token and total memory at your target concurrency.

- Weight, activation and KV cache breakdown per model
- Quantization support (FP16, INT8, INT4, GPTQ, AWQ)
- Concurrency and batch-size aware
- Recommends GPU class and count

## Why it matters

Tokenomy is the economic runtime for AI — it shows where every AI dollar goes and gives you the controls to redirect it. This tool is free to use; sign in to save scenarios, share with your team and connect it to your live usage ledger.

## Related pages

- [Token Calculator](https://tokenomy.ai/tools/token-calculator)
- [Token Observability](https://tokenomy.ai/tools/token-observability)
- [Token Speed Simulator](https://tokenomy.ai/tools/token-speed-simulator)
- [Memory Calculator](https://tokenomy.ai/tools/memory-calculator)
- [Energy Usage Estimator](https://tokenomy.ai/tools/energy-usage-estimator)
- [Token Leaderboard](https://tokenomy.ai/tools/token-leaderboard)
- [Alternatives Explorer](https://tokenomy.ai/tools/alternatives-explorer)
- [Cost Optimization Suite](https://tokenomy.ai/tools/cost-optimization)
- [Why FinOps for AI](https://tokenomy.ai/finops-for-ai)
- [Pricing](https://tokenomy.ai/pricing)
