Skip to content

Llama 3.3 70B Instruct

Metatext → textHugging Face OpenRouter

The cheapest provider for Llama 3.3 70B Instruct is DeepInfra at $0.155 per 1M tokens (blended), charging $0.10 per 1M input and $0.32 per 1M output tokens. 10 providers serve it; the most expensive charges 6.7× as much. Context window: 128K tokens. Prices tracked since Oct 3, 2026.

Best price
$0.155
at DeepInfra
Input / output
$0.10 / $0.32
per 1M tokens
Providers
10
6.7× spread
7 days
—
best price
30 days
—
best price
Context
128K
tokens

Providers

Active offers ranked by blended price. Bars are relative to the most expensive.

  • DeepInfraCheapest
    turbo
    $0.155
    $0.10 in · $0.32 out
    best
    128K ctx98.5% uptimefp8
  • $0.201
    $0.135 in · $0.40 out
    +30%
    12K ctx89.5% uptimebf16
  • $0.28
    $0.20 in · $0.52 out
    +81%
    128K ctx99.5% uptimefp8
  • $0.29
    $0.22 in · $0.50 out
    +87%
    128K ctx99.3% uptimefp8
  • $0.563
    $0.45 in · $0.90 out
    +263%
    128K ctx98.6% uptime−25% discount
  • $0.64
    $0.59 in · $0.79 out
    +313%
    128K ctx99.8% uptime
  • $0.71
    $0.71 in · $0.71 out
    +358%
    125K ctx99.8% uptimefp16
  • $0.72
    $0.72 in · $0.72 out
    +365%
    125K ctx
  • us-central1
    $0.72
    $0.72 in · $0.72 out
    +365%
    125K ctx
  • $0.783
    $0.293 in · $2.253 out
    +405%
    24K ctx99.1% uptimefp8
  • $1.04
    $1.04 in · $1.04 out
    +571%
    128K ctx98.0% uptime
Blended = (3 × input + output) / 4.

Price history

Hover to compare providers at any point in time. Click a legend item to hide it.

USD per 1M tokens · each step is a price change
Not enough history yet
Tracking since Oct 3, 2026. The chart fills in as new crawls arrive every 6 hours.

Other pricing

Cache, reasoning and per-call fees. Token rates per 1M; others per unit.

Offered byCache read
AkashML$0.10
Parasail$0.11
Groq$0.295
CoreWeave$0.71

Change log

Every recorded change for this model since Oct 3, 2026.

No changes yet
Tracking since Oct 3, 2026. Price changes, new providers and delistings will show up here.

About this model

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Released
Dec 6, 2024
First tracked
Oct 3, 2026
Canonical slug
meta-llama/llama-3.3-70b-instruct
Weights
meta-llama/Llama-3.3-70B-Instruct