GLM 5.3 Flash
The cheapest provider for GLM 5.3 Flash is DeepInfra at $0.119 per 1M tokens (blended), charging $0.075 per 1M input and $0.25 per 1M output tokens. 32 providers serve it; the most expensive charges 4.0× as much. Context window: 1M tokens. Prices tracked since Oct 3, 2026.
Best price
$0.119
at DeepInfra
Input / output
$0.075 / $0.25
per 1M tokens
Providers
32
4.0× spread
7 days
—
best price
30 days
—
best price
Context
1M
tokens
Providers
Active offers ranked by blended price. Bars are relative to the most expensive.
| Provider | Input | Output | Blended | vs cheapest | Context | Uptime 1d | Price since |
|---|---|---|---|---|---|---|---|
DeepInfraCheapest fp4·−50% discount | $0.075 | $0.25 | $0.119 | best | |||
| fp8·−44% discount | $0.0839 | $0.28 | $0.133 | +12% | |||
| fp8·−44% discount | $0.084 | $0.28 | $0.133 | +12% | |||
| fp8·−40% discount | $0.09 | $0.30 | $0.143 | +20% | |||
| $0.035 | $0.50 | $0.151 | +27% | ||||
| fp8·−30% discount | $0.105 | $0.35 | $0.166 | +40% | |||
| $0.10 | $0.40 | $0.175 | +47% | ||||
| fp4 | $0.042 | $0.60 | $0.182 | +53% | |||
| fp4 | $0.045 | $0.60 | $0.184 | +55% | |||
| fp8·−20% discount | $0.12 | $0.40 | $0.19 | +60% | |||
| fp4 | $0.0455 | $0.65 | $0.197 | +66% | |||
| fp4·−15% discount | $0.128 | $0.425 | $0.202 | +70% | |||
| fp8·−10% discount | $0.135 | $0.45 | $0.214 | +80% | |||
| fp8 | $0.15 | $0.50 | $0.238 | +100% | |||
| fp8 | $0.15 | $0.50 | $0.238 | +100% | |||
| nvfp4 | $0.15 | $0.50 | $0.238 | +100% | |||
| fp4 | $0.15 | $0.50 | $0.238 | +100% | |||
| $0.15 | $0.50 | $0.238 | +100% | ||||
| $0.15 | $0.50 | $0.238 | +100% | ||||
| $0.15 | $0.50 | $0.238 | +100% | ||||
| nvfp4 | $0.15 | $0.50 | $0.238 | +100% | |||
| fp4 | $0.15 | $0.50 | $0.238 | +100% | |||
| $0.15 | $0.50 | $0.238 | +100% | ||||
| fp8 | $0.15 | $0.50 | $0.238 | +100% | |||
| $0.15 | $0.50 | $0.238 | +100% | ||||
| $0.15 | $0.50 | $0.238 | +100% | ||||
| fp8 | $0.15 | $0.50 | $0.238 | +100% | |||
| fp8 | $0.17 | $0.48 | $0.248 | +108% | |||
| fp8 | $0.165 | $0.55 | $0.261 | +120% | |||
| fp8 | $0.225 | $0.45 | $0.281 | +137% | |||
| $0.225 | $0.75 | $0.356 | +200% | ||||
| $0.10 | $1.25 | $0.388 | +226% | ||||
| $0.30 | $1.00 | $0.475 | +300% |
- DeepInfraCheapest$0.119$0.075 in · $0.25 outbest1M ctx99.8% uptimefp4·−50% discount
- $0.133$0.0839 in · $0.28 out+12%1M ctx97.0% uptimefp8·−44% discount
- $0.133$0.084 in · $0.28 out+12%1M ctx98.2% uptimefp8·−44% discount
- $0.143$0.09 in · $0.30 out+20%1M ctx99.0% uptimefp8·−40% discount
- $0.151$0.035 in · $0.50 out+27%1M ctx99.6% uptime
- $0.166$0.105 in · $0.35 out+40%1M ctx97.9% uptimefp8·−30% discount
- $0.175$0.10 in · $0.40 out+47%1M ctx99.2% uptime
- $0.182$0.042 in · $0.60 out+53%1M ctx99.5% uptimefp4
- $0.184$0.045 in · $0.60 out+55%1M ctx99.3% uptimefp4
- $0.19$0.12 in · $0.40 out+60%1M ctx99.0% uptimefp8·−20% discount
- $0.197$0.0455 in · $0.65 out+66%1M ctx99.6% uptimefp4
- $0.202$0.128 in · $0.425 out+70%1M ctx99.8% uptimefp4·−15% discount
- $0.214$0.135 in · $0.45 out+80%262K ctx99.0% uptimefp8·−10% discount
- $0.238$0.15 in · $0.50 out+100%1M ctx91.7% uptimefp8
- $0.238$0.15 in · $0.50 out+100%1M ctx100% uptimefp8
- $0.238$0.15 in · $0.50 out+100%1M ctx99.9% uptimenvfp4
- $0.238$0.15 in · $0.50 out+100%1M ctx98.7% uptimefp4
- $0.238$0.15 in · $0.50 out+100%1M ctx98.7% uptime
- $0.238$0.15 in · $0.50 out+100%1M ctx99.3% uptime
- $0.238$0.15 in · $0.50 out+100%1M ctx99.5% uptime
- $0.238$0.15 in · $0.50 out+100%1M ctx99.8% uptimenvfp4
- $0.238$0.15 in · $0.50 out+100%1M ctx99.1% uptimefp4
- $0.238$0.15 in · $0.50 out+100%256K ctx99.8% uptime
- $0.238$0.15 in · $0.50 out+100%1M ctx98.2% uptimefp8
- $0.238$0.15 in · $0.50 out+100%1M ctx99.8% uptime
- $0.238$0.15 in · $0.50 out+100%1M ctx97.9% uptime
- $0.238$0.15 in · $0.50 out+100%1M ctx99.8% uptimefp8
- $0.248$0.17 in · $0.48 out+108%1M ctx99.1% uptimefp8
- $0.261$0.165 in · $0.55 out+120%1M ctx99.9% uptimefp8
- $0.281$0.225 in · $0.45 out+137%1M ctx98.5% uptimefp8
- $0.356$0.225 in · $0.75 out+200%1M ctx99.8% uptime
- $0.388$0.10 in · $1.25 out+226%1M ctx99.8% uptime
- $0.475$0.30 in · $1.00 out+300%1M ctx100% uptime
Blended = (3 × input + output) / 4.
Price history
Hover to compare providers at any point in time. Click a legend item to hide it.
USD per 1M tokens · each step is a price change
Not enough history yet
Tracking since Oct 3, 2026. The chart fills in as new crawls arrive every 6 hours.
Other pricing
Cache, reasoning and per-call fees. Token rates per 1M; others per unit.
| Offered by | Cache read |
|---|---|
| DeepInfra | $0.015 |
| StreamLake | $0.0168 |
| Novita | $0.0168 |
| GMICloud | $0.018 |
| Relace | $0.035 |
| Near AI | $0.0245 |
| DekaLLM, InferenceNet +1 | $0.04 |
| Sail Research (us) | $0.0285 |
| Phala | $0.024 |
| OpenInference | $0.0455 |
| Decart | $0.0255 |
| Io Net | $0.027 |
| AtlasCloud, BaseTen +12 | $0.03 |
| CoreWeave | $0.05 |
| NextBit | $0.033 |
| Inceptron, Wafer | $0.08 |
| Fireworks (us) | $0.045 |
Change log
Every recorded change for this model since Oct 3, 2026.
No changes yet
Tracking since Oct 3, 2026. Price changes, new providers and delistings will show up here.
About this model
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- Released
- Aug 26, 2026
- First tracked
- Oct 3, 2026
- Canonical slug
- z-ai/glm-5.3-flash-20260826
- Weights
- zai-org/GLM-5.3-Flash