llama-3.3-70b pricing

llama-3.3-70b is priced at $0.60 per 1M input tokens and $0.60 per 1M output tokens, with roughly 350 ms to first token and about 45 tokens/sec output throughput.

Price per 1M tokens

Input $0.60 · Output $0.60 · Output/input ratio 1.0x.

  • 1M input + 1M output tokens: $1.20
  • 10M input + 2M output tokens/month: $7.20
  • 100M input + 20M output tokens/month: $72.00

Cheaper alternatives to llama-3.3-70b

Same workload, lower output price. Validate quality before switching.

  • gemini-3.5-flash — $0.50/1M output (17% cheaper)
  • llama-4-maverick — $0.50/1M output (17% cheaper)
  • grok-3-mini — $0.50/1M output (17% cheaper)
  • gemini-3-flash — $0.40/1M output (33% cheaper)

How much are you wasting on llama-3.3-70b?

Price lists are the easy part. Run the free waste scanner on your usage export to see which calls should have run on a cheaper tier, which prompts are oversized, and which retries you paid for twice — with dollars recoverable per finding.