Skip to main content
DeepSeek V4.1 Flash is available through AnyFast with the public model ID deepseek-v4.1-flash and the OpenAI-compatible POST /v1/chat/completions endpoint. It supports text and image input, thinking and non-thinking modes, function calling, JSON output, streaming, and prompt caching.
DeepSeek calls this release DeepSeek-V4.1-Flash and uses deepseek-flash on its own API. When calling AnyFast, use deepseek-v4.1-flash.

Model specifications

Quick examples

cURL

Thinking mode

Thinking mode is enabled by default with high effort. Use {"thinking":{"type":"disabled"}} for latency-sensitive requests that do not need reasoning. When thinking is enabled, reasoning is returned in choices[0].message.reasoning_content and the final answer is returned in choices[0].message.content. The official model accepts low, high, and max effort. temperature, presence_penalty, and frequency_penalty have no effect in thinking mode. If a request includes tools, pass the complete reasoning_content from earlier assistant messages back in subsequent tool-call turns.

Image input

Use image_url blocks only in user messages. Supported provider formats are JPEG, PNG, GIF, and WebP. For URL inputs, use a directly downloadable HTTP(S) address; for local images, use a Base64 data URL. Image tokens are included in usage.prompt_tokens.

Common parameters

DeepSeek V4.1 Flash API Reference

View the complete Chat Completions request and response schema.
Sources: DeepSeek V4.1 Flash release, models and pricing, thinking mode, and vision.