Meta

Llama 3.3 70B Pricing

Llama 3.3 70B is a budget-tier language model from Meta, priced at $0.12 per 1M input tokens and $0.30 per 1M output tokens. Prices via hosted providers (DeepInfra, Groq, etc.).

Input
$0.12
per 1M tokens
Output
$0.30
per 1M tokens
Context
128K
tokens
Max output
16K
tokens

What does Llama 3.3 70B cost in practice?

These are real-world monthly estimates based on common workload sizes. They assume a 2:1 input-to-output ratio (typical of chat + retrieval-augmented apps).

Workload Input / mo Output / mo Monthly cost Annualized
Light prototyping 0.1M tokens 0.1M tokens $0.03 $0.32
Medium production 5.0M tokens 2.5M tokens $1.35 $16.20
Heavy production 50.0M tokens 25.0M tokens $13.50 $162.00

Need to model your own usage? Use the LLM pricing calculator →

Llama 3.3 70B features

Alternatives to Llama 3.3 70B

Cheaper, same-tier alternatives

If Llama 3.3 70B is more than you need, these budget-tier models cost less per token:

Higher-priced same-tier models

If you've outgrown Llama 3.3 70B, consider these:

Need the cheapest Meta option? Llama 4 Scout starts at $ 0.15 input / $ 0.50 output per 1M tokens →

FAQ

How much does Llama 3.3 70B cost per 1M tokens?

Llama 3.3 70B costs $0.12 per 1M input tokens and $0.30 per 1M output tokens.

What is Llama 3.3 70B's context window?

Llama 3.3 70B supports a 128,000-token context window with a maximum output of 16,000 tokens per response.

What features does Llama 3.3 70B support?

Tool Use, Streaming.

How much does Llama 3.3 70B cost for medium production usage?

For ~5M input + 2.5M output tokens per month, Llama 3.3 70B costs approximately $1.35/mo (≈$ 16.20/yr).

What are cheaper alternatives to Llama 3.3 70B?

Same-tier alternatives that cost less: Mistral Small 3.1 .