# llama-4.1-405b pricing

> llama-4.1-405b costs $2.80 per 1M input tokens and $8.00 per 1M output tokens. Compare against cheaper alternatives and estimate your monthly bill — free, no login.

Source: https://tokenomy.ai/models/llama-4-1-405b
Last updated: 2026-09-12
Publisher: Tokenomy — FinOps for AI
License: free to quote with attribution and a link to the source URL.

## Answer

llama-4.1-405b costs $2.80 per 1M input tokens and $8.00 per 1M output tokens as of 2026-09-12. A workload of 10M input and 2M output tokens per month runs $44.00. The cheapest comparable alternative is claude-haiku-4.5 at $2.50 per 1M output tokens.

llama-4.1-405b is priced at $2.80 per 1M input tokens and $8.00 per 1M output tokens, with roughly 320 ms to first token and about 60 tokens/sec output throughput.

## Price per 1M tokens

Input $2.80 · Output $8.00 · Output/input ratio 2.9x.

- 1M input + 1M output tokens: $10.80
- 10M input + 2M output tokens/month: $44.00
- 100M input + 20M output tokens/month: $440.00

## Cheaper alternatives to llama-4.1-405b

Same workload, lower output price. Validate quality before switching.

- mistral-large-4 — $6.00/1M output (25% cheaper)
- gpt-5-mini — $3.20/1M output (60% cheaper)
- amazon-nova-pro — $3.20/1M output (60% cheaper)
- claude-haiku-4.5 — $2.50/1M output (69% cheaper)

## How much are you wasting on llama-4.1-405b?

Price lists are the easy part. Run the free waste scanner on your usage export to see which calls should have run on a cheaper tier, which prompts are oversized, and which retries you paid for twice — with dollars recoverable per finding.

## Frequently asked questions

### How much does llama-4.1-405b cost per million tokens?

$2.80 per 1M input tokens and $8.00 per 1M output tokens, as listed on 2026-09-12.

### What does llama-4.1-405b cost per month at typical usage?

At 10M input and 2M output tokens per month, llama-4.1-405b costs about $44.00. At 100M input and 20M output tokens it costs about $440.00.

### Is there a cheaper alternative to llama-4.1-405b?

Yes. mistral-large-4 at $6.00/1M output, gpt-5-mini at $3.20/1M output, amazon-nova-pro at $3.20/1M output, claude-haiku-4.5 at $2.50/1M output. Validate quality on your own evaluation set before switching.

### How fast is llama-4.1-405b?

Roughly 320 ms to first token and about 60 output tokens per second.

## Related pages

- [Scan your AI spend free](https://tokenomy.ai/waste)
- [All model pricing](https://tokenomy.ai/models)
- [mistral-large-4 pricing](https://tokenomy.ai/models/mistral-large-4)
- [gpt-5-mini pricing](https://tokenomy.ai/models/gpt-5-mini)
- [amazon-nova-pro pricing](https://tokenomy.ai/models/amazon-nova-pro)
- [Token cost calculator](https://tokenomy.ai/tools/token-calculator)
