llama-4-scout pricing

llama-4-scout costs $0.20 per 1M input tokens and $0.20 per 1M output tokens as of 2026-09-12. A workload of 10M input and 2M output tokens per month runs $2.40. The cheapest comparable alternative is gemini-2.5-flash-lite at $0.15 per 1M output tokens.

llama-4-scout is priced at $0.20 per 1M input tokens and $0.20 per 1M output tokens, with roughly 180 ms to first token and about 70 tokens/sec output throughput.

Last updated . Model pricing is refreshed twice daily.

Price per 1M tokens

Input $0.20 · Output $0.20 · Output/input ratio 1.0x.

  • 1M input + 1M output tokens: $0.40
  • 10M input + 2M output tokens/month: $2.40
  • 100M input + 20M output tokens/month: $24.00

Cheaper alternatives to llama-4-scout

Same workload, lower output price. Validate quality before switching.

  • gemini-2.5-flash-lite — $0.15/1M output (25% cheaper)

How much are you wasting on llama-4-scout?

Price lists are the easy part. Run the free waste scanner on your usage export to see which calls should have run on a cheaper tier, which prompts are oversized, and which retries you paid for twice — with dollars recoverable per finding.

Frequently asked questions

How much does llama-4-scout cost per million tokens?

$0.20 per 1M input tokens and $0.20 per 1M output tokens, as listed on 2026-09-12.

What does llama-4-scout cost per month at typical usage?

At 10M input and 2M output tokens per month, llama-4-scout costs about $2.40. At 100M input and 20M output tokens it costs about $24.00.

Is there a cheaper alternative to llama-4-scout?

Yes. gemini-2.5-flash-lite at $0.15/1M output. Validate quality on your own evaluation set before switching.

How fast is llama-4-scout?

Roughly 180 ms to first token and about 70 output tokens per second.