Qwen3.8 2.4T A95B

  • Tool calling
  • Streaming
  • Open weights

Model id

qwen/qwen3.8-2.4t-a95b
Context window
256K
tokens
Input price
$2
per 1M tokens
Output price
$6
per 1M tokens
Max output
Provider
no narrower limit

Provider list price

Pricing

USD / 1M tokens

Input$2 / M tokens
Output$6 / M tokens
Cache read$0.2 / M tokens
Context window256K tokens
RoutingFree

Provider list price. BYOK bills your provider; managed credits draw down at these rates. No markup or request fee. Every request is metered per engineer and joined to shipped work.

Registry 2026-09-29.1 · effective 9/17/2026 · Private rate card. Source withheld.

Registry contract

Capabilities

Protocols
3
Chat · Responses · Messages · streaming
Tools
2
Chat · Messages
JSON Schema
None
Not supported
Vision
No
Not supported
Reasoning
None
No public control
Output limit
Provider
No narrower TokenSpend limit

canonical qwen/qwen3.8-2.4t-a95b · revision qwen/qwen3.8-2.4t-a95b · aliases qwen3.8-2.4t-a95b

First request

Run it

Python · TokenSpend SDK
from tokenspend import TokenSpend

client = TokenSpend()
resp = client.chat.completions.create(
    model="qwen/qwen3.8-2.4t-a95b",
    messages=[{"role": "user", "content": "hello"}],
)
TypeScript · TokenSpend SDK
import { TokenSpend } from "@tokenspend/sdk";

const client = new TokenSpend();
const resp = await client.chat.completions.create({
  model: "qwen/qwen3.8-2.4t-a95b",
  messages: [{ role: "user", content: "hello" }]
});
curl
curl https://api.tokenspend.dev/v1/chat/completions \
  -H "content-type: application/json" \
  -H "Authorization: Bearer $TOKENSPEND_API_KEY" \
  -d '{"model":"qwen/qwen3.8-2.4t-a95b","messages":[{"role":"user","content":"hello"}]}'
Claude Code integration
export ANTHROPIC_BASE_URL=https://api.tokenspend.dev
export ANTHROPIC_API_KEY=$TOKENSPEND_API_KEY
export ANTHROPIC_MODEL=qwen/qwen3.8-2.4t-a95b

One request shape reaches every model. Claude Code connects as an integration with two environment variables. TokenSpend never stores prompts or responses.

Questions

How much does Qwen3.8 2.4T A95B cost?

$2 per 1M input tokens and $6 per 1M output tokens, at provider list price. TokenSpend adds no markup or request fee. Managed credits carry a 5% fee when you buy them.

What is the context window of Qwen3.8 2.4T A95B?

Qwen3.8 2.4T A95B has a 256K token context window.

Does Qwen3.8 2.4T A95B support tool calling?

Yes. Tool calling is supported on Chat, Messages.

Does Qwen3.8 2.4T A95B support images?

No. Image input is not supported on Qwen3.8 2.4T A95B through TokenSpend.

Does Qwen3.8 2.4T A95B support reasoning?

No. Reasoning control is not supported on Qwen3.8 2.4T A95B through TokenSpend.

Does Qwen3.8 2.4T A95B have prompt caching?

Yes. Cache reads cost $0.2 per 1M tokens.

When were these prices last checked?

Prices follow rate card 2026-09-29.1, effective 9/17/2026. Source: Private rate card.