All models
Z.ai
GLM-5.3 Flash
open weightsimages
- Input
- $0.075
- per M tokens
- Output
- $0.25
- per M tokens
- Cache read
- $0.015
- per M tokens
- Context
- 1,048.576K
- tokens
Registry contract
Capabilities
zai/glm-5.3-flash- Protocols
- 3
- Chat · Responses · Messages · streaming
- Tools
- None
- Not supported
- JSON Schema
- None
- Not supported
- Vision
- Yes
- Chat · Responses · Messages
- Reasoning
- None
- No public control
- Output limit
- Provider
- No narrower TokenSpend limit
canonical zai/glm-5.3-flash · revision zai/glm-5.3-flash · aliases glm-5.3-flash, zai-org/GLM-5.3-Flash
Provider list price
Pricing
USD / 1M tokens
| Input | $0.075 / M tokens |
| Output | $0.25 / M tokens |
| Cache read | $0.015 / M tokens |
| Routing | Free |
Provider list price. BYOK bills your provider; managed credits draw down at these rates. No markup or request fee. Every request is metered per engineer and joined to shipped work.
Registry 2026-09-15.1 · effective 9/15/2026 · Private rate card. Source withheld.
First request
Run it
Python · TokenSpend SDK
from tokenspend import TokenSpend
client = TokenSpend()
resp = client.chat.completions.create(
model="zai/glm-5.3-flash",
messages=[{"role": "user", "content": "hello"}],
)TypeScript · TokenSpend SDK
import { TokenSpend } from "@tokenspend/sdk";
const client = new TokenSpend();
const resp = await client.chat.completions.create({
model: "zai/glm-5.3-flash",
messages: [{ role: "user", content: "hello" }]
});curl
curl https://api.tokenspend.dev/v1/chat/completions \
-H "content-type: application/json" \
-H "Authorization: Bearer $TOKENSPEND_API_KEY" \
-d '{"model":"zai/glm-5.3-flash","messages":[{"role":"user","content":"hello"}]}'Claude Code integration
export ANTHROPIC_BASE_URL=https://api.tokenspend.dev export ANTHROPIC_API_KEY=$TOKENSPEND_API_KEY export ANTHROPIC_MODEL=zai/glm-5.3-flash
One request shape reaches every model. Claude Code connects as an integration with two environment variables. Prompts are not stored unless you enable request logging.