llama-4-behemoth pricing
llama-4-behemoth costs $1.00 per 1M input tokens and $1.00 per 1M output tokens as of 2026-09-12. A workload of 10M input and 2M output tokens per month runs $12.00. The cheapest comparable alternative is gemini-2.5-flash at $0.60 per 1M output tokens.
llama-4-behemoth is priced at $1.00 per 1M input tokens and $1.00 per 1M output tokens, with roughly 400 ms to first token and about 40 tokens/sec output throughput.
Last updated . Model pricing is refreshed twice daily.
Price per 1M tokens
Input $1.00 · Output $1.00 · Output/input ratio 1.0x.
- 1M input + 1M output tokens: $2.00
- 10M input + 2M output tokens/month: $12.00
- 100M input + 20M output tokens/month: $120.00
Cheaper alternatives to llama-4-behemoth
Same workload, lower output price. Validate quality before switching.
- amazon-nova-lite — $0.80/1M output (20% cheaper)
- mistral-small-3 — $0.80/1M output (20% cheaper)
- llama-3.3-70b — $0.60/1M output (40% cheaper)
- gemini-2.5-flash — $0.60/1M output (40% cheaper)
How much are you wasting on llama-4-behemoth?
Price lists are the easy part. Run the free waste scanner on your usage export to see which calls should have run on a cheaper tier, which prompts are oversized, and which retries you paid for twice — with dollars recoverable per finding.
Frequently asked questions
How much does llama-4-behemoth cost per million tokens?
$1.00 per 1M input tokens and $1.00 per 1M output tokens, as listed on 2026-09-12.
What does llama-4-behemoth cost per month at typical usage?
At 10M input and 2M output tokens per month, llama-4-behemoth costs about $12.00. At 100M input and 20M output tokens it costs about $120.00.
Is there a cheaper alternative to llama-4-behemoth?
Yes. amazon-nova-lite at $0.80/1M output, mistral-small-3 at $0.80/1M output, llama-3.3-70b at $0.60/1M output, gemini-2.5-flash at $0.60/1M output. Validate quality on your own evaluation set before switching.
How fast is llama-4-behemoth?
Roughly 400 ms to first token and about 40 output tokens per second.