DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp. It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

Input price
$0.30USD / 1M tokens
Output price
$1.20USD / 1M tokens
Context window
N/A
Max output
943.7K

Try the model

Playground

Open playground
Input⌘ / Ctrl + Enter

Try a prompt

943.7K max output
Input $0.30/M · Output $1.20/M · Cache read $0.0060/M
OutputReady

Your response will appear here

Choose an example or write a prompt, then click Run.

Specifications

Pricing

Input
$0.30 / 1M tokens
Output
$1.20 / 1M tokens
Billing formula
usage.prompt_tokens * 0.3 / 1000000 + usage.completion_tokens * 1.2 / 1000000 + usage.cached_tokens * 0.000006 / 1000000

Context & modalities

Max output
943,718 tokens
Input
Text
Output
Text

Capabilities

Chat
Supported
Vision
Not supported
Reasoning
Supported
Structured output
Supported
Function calling
Supported
Audio input
Not supported

Access

Provider
DeepSeek
Model ID
deepseek/deepseek-v4.1-flash
Execution
sync
API
Chat Completions API
Endpoint
/v1/chat/completions

Start building

Send your first request

OpenAI-compatible endpoint with unified authentication and usage tracking.

Production API
Chat Completions API endpoint
https://api.sandbase.ai/v1/chat/completions
Model ID
deepseek/deepseek-v4.1-flash
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
  -X POST "https://api.sandbase.ai/v1/chat/completions" \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'SANDBASE_JSON'
{
  "model": "deepseek/deepseek-v4.1-flash",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ],
  "max_tokens": 2048
}
SANDBASE_JSON
)
printf '%s\n' "$result"

Choose your model

Compare models

All language models
ModelContextInput / 1MOutput / 1MReleased
DeepSeek—$0.30$1.20Sep 10, 2026
DeepSeek1M——Aug 21, 2026
DeepSeek1M——Jul 31, 2026
DeepSeek1M——Apr 24, 2026
DeepSeek1M——Apr 24, 2026
DeepSeek160K——Dec 1, 2025