DeepSeek V4 Flash 0731

  • Tool calling
  • JSON schema
  • Reasoning
  • Streaming
  • 1.05M context
  • Open weights

Model id

deepseek/deepseek-v4-flash-0731
Context window
1.05M
tokens
Input price
$0.06
per 1M tokens
Output price
$0.18
per 1M tokens
Max output
Provider
no narrower limit

Provider list price

Pricing

USD / 1M tokens

Input$0.06 / M tokens
Output$0.18 / M tokens
Cache read$0.015 / M tokens
Context window1.05M tokens
RoutingFree

Provider list price. BYOK bills your provider; managed credits draw down at these rates. No markup or request fee. Every request is metered per engineer and joined to shipped work.

Registry 2026-09-29.1 · effective 9/17/2026 · Private rate card. Source withheld.

Registry contract

Capabilities

Protocols
3
Chat · Responses · Messages · streaming
Tools
1
Chat
JSON Schema
2
Chat · Responses
Vision
No
Not supported
Reasoning
1
Chat
Output limit
Provider
No narrower TokenSpend limit

canonical deepseek/deepseek-v4-flash-0731 · revision deepseek/deepseek-v4-flash-0731 · aliases deepseek-ai/DeepSeek-V4-Flash-0731, deepseek-v4-flash-0731

First request

Run it

Python · TokenSpend SDK
from tokenspend import TokenSpend

client = TokenSpend()
resp = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash-0731",
    messages=[{"role": "user", "content": "hello"}],
)
TypeScript · TokenSpend SDK
import { TokenSpend } from "@tokenspend/sdk";

const client = new TokenSpend();
const resp = await client.chat.completions.create({
  model: "deepseek/deepseek-v4-flash-0731",
  messages: [{ role: "user", content: "hello" }]
});
curl
curl https://api.tokenspend.dev/v1/chat/completions \
  -H "content-type: application/json" \
  -H "Authorization: Bearer $TOKENSPEND_API_KEY" \
  -d '{"model":"deepseek/deepseek-v4-flash-0731","messages":[{"role":"user","content":"hello"}]}'
Claude Code integration
export ANTHROPIC_BASE_URL=https://api.tokenspend.dev
export ANTHROPIC_API_KEY=$TOKENSPEND_API_KEY
export ANTHROPIC_MODEL=deepseek/deepseek-v4-flash-0731

One request shape reaches every model. Claude Code connects as an integration with two environment variables. TokenSpend never stores prompts or responses.

Questions

How much does DeepSeek V4 Flash 0731 cost?

$0.06 per 1M input tokens and $0.18 per 1M output tokens, at provider list price. TokenSpend adds no markup or request fee. Managed credits carry a 5% fee when you buy them.

What is the context window of DeepSeek V4 Flash 0731?

DeepSeek V4 Flash 0731 has a 1.05M token context window.

Does DeepSeek V4 Flash 0731 support tool calling?

Yes. Tool calling is supported on Chat.

Does DeepSeek V4 Flash 0731 support images?

No. Image input is not supported on DeepSeek V4 Flash 0731 through TokenSpend.

Does DeepSeek V4 Flash 0731 support reasoning?

Yes. Reasoning control is supported on Chat.

Does DeepSeek V4 Flash 0731 have prompt caching?

Yes. Cache reads cost $0.015 per 1M tokens.

When were these prices last checked?

Prices follow rate card 2026-09-29.1, effective 9/17/2026. Source: Private rate card.