deepseek-v4.1-flash and the OpenAI-compatible POST /v1/chat/completions endpoint. It supports text and image input, thinking and non-thinking modes, function calling, JSON output, streaming, and prompt caching.
DeepSeek calls this release
DeepSeek-V4.1-Flash and uses deepseek-flash on its own API. When calling AnyFast, use deepseek-v4.1-flash.Model specifications
Quick examples
- Text
- Vision
- Thinking
cURL
Thinking mode
Thinking mode is enabled by default withhigh effort. Use {"thinking":{"type":"disabled"}} for latency-sensitive requests that do not need reasoning. When thinking is enabled, reasoning is returned in choices[0].message.reasoning_content and the final answer is returned in choices[0].message.content.
The official model accepts low, high, and max effort. temperature, presence_penalty, and frequency_penalty have no effect in thinking mode. If a request includes tools, pass the complete reasoning_content from earlier assistant messages back in subsequent tool-call turns.
Image input
Useimage_url blocks only in user messages. Supported provider formats are JPEG, PNG, GIF, and WebP. For URL inputs, use a directly downloadable HTTP(S) address; for local images, use a Base64 data URL. Image tokens are included in usage.prompt_tokens.
Common parameters
DeepSeek V4.1 Flash API Reference
View the complete Chat Completions request and response schema.