> ## Documentation Index
> Fetch the complete documentation index at: https://docs.anyfast.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# DeepSeek V4.1 Flash

> Use DeepSeek V4.1 Flash through the AnyFast OpenAI-compatible Chat Completions API.

DeepSeek V4.1 Flash is available through AnyFast with the public model ID `deepseek-v4.1-flash` and the OpenAI-compatible `POST /v1/chat/completions` endpoint. It supports text and image input, thinking and non-thinking modes, function calling, JSON output, streaming, and prompt caching.

<Info>
  DeepSeek calls this release `DeepSeek-V4.1-Flash` and uses `deepseek-flash` on its own API. When calling AnyFast, use `deepseek-v4.1-flash`.
</Info>

## Model specifications

| Property | Value |
| - | - |
| AnyFast model ID | `deepseek-v4.1-flash` |
| Input | Text and images |
| Output | Text |
| Context window | 1,048,576 tokens |
| Maximum output | 393,216 tokens |
| Thinking mode | Supported and enabled by default |
| Thinking effort | `low`, `high`, `max`; default `high` |
| Structured output | JSON mode |
| Tools | Function calling |
| Streaming | SSE |

## Quick examples

<Tabs sync={false}>
  <Tab title="Text">
    ```bash cURL theme={null}
    curl https://www.anyfast.ai/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "deepseek-v4.1-flash",
        "messages": [
          {"role": "user", "content": "Explain how a B-tree index accelerates database queries."}
        ],
        "max_tokens": 1024
      }'
    ```
  </Tab>

  <Tab title="Vision">
    Pass the user message as an array containing text and `image_url` blocks. The image can use a public HTTP(S) URL or a Base64 data URL.

    ```bash cURL theme={null}
    curl https://www.anyfast.ai/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "deepseek-v4.1-flash",
        "messages": [{
          "role": "user",
          "content": [
            {"type": "text", "text": "Describe this image and read any visible text."},
            {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}}
          ]
        }]
      }'
    ```
  </Tab>

  <Tab title="Thinking">
    `thinking` controls whether reasoning is enabled. Set `reasoning_effort` separately at the top level.

    ```python Python theme={null}
    from openai import OpenAI

    client = OpenAI(
        api_key="YOUR_API_KEY",
        base_url="https://www.anyfast.ai/v1",
    )

    response = client.chat.completions.create(
        model="deepseek-v4.1-flash",
        messages=[
            {"role": "user", "content": "Find and explain the bug in this algorithm."}
        ],
        reasoning_effort="high",
        extra_body={"thinking": {"type": "enabled"}},
    )

    print(response.choices[0].message.reasoning_content)
    print(response.choices[0].message.content)
    ```
  </Tab>
</Tabs>

## Thinking mode

Thinking mode is enabled by default with `high` effort. Use `{"thinking":{"type":"disabled"}}` for latency-sensitive requests that do not need reasoning. When thinking is enabled, reasoning is returned in `choices[0].message.reasoning_content` and the final answer is returned in `choices[0].message.content`.

The official model accepts `low`, `high`, and `max` effort. `temperature`, `presence_penalty`, and `frequency_penalty` have no effect in thinking mode. If a request includes `tools`, pass the complete `reasoning_content` from earlier assistant messages back in subsequent tool-call turns.

## Image input

Use `image_url` blocks only in `user` messages. Supported provider formats are JPEG, PNG, GIF, and WebP. For URL inputs, use a directly downloadable HTTP(S) address; for local images, use a Base64 data URL. Image tokens are included in `usage.prompt_tokens`.

## Common parameters

| Parameter | Type | Required | Description |
| - | - | -: | - |
| `model` | string | Yes | Must be `deepseek-v4.1-flash`. |
| `messages` | array | Yes | Conversation messages using `system`, `user`, `assistant`, or `tool` roles. |
| `thinking` | object | No | `{"type":"enabled"}` or `{"type":"disabled"}`. Thinking is enabled by default. |
| `reasoning_effort` | string | No | `none` disables thinking; `low`, `high`, and `max` select the effort. Default: `high`. |
| `max_tokens` | integer | No | Maximum generated tokens, from 1 through 393,216. |
| `stream` | boolean | No | Stream Chat Completions chunks over SSE. |
| `response_format` | object | No | Use `{"type":"json_object"}` for JSON mode and explicitly ask for JSON in the prompt. |
| `tools` | array | No | OpenAI-format function tools. |
| `tool_choice` | string or object | No | `none`, `auto`, `required`, or a named function. |
| `user_id` | string | No | Stable end-user identifier for content safety and cache isolation. |

<Card title="DeepSeek V4.1 Flash API Reference" icon="code" href="/api-reference/model-api/deepseek/deepseek-v4-1-flash">
  View the complete Chat Completions request and response schema.
</Card>

Sources: [DeepSeek V4.1 Flash release](https://api-docs.deepseek.com/news/news260910/), [models and pricing](https://api-docs.deepseek.com/quick_start/pricing/), [thinking mode](https://api-docs.deepseek.com/guides/thinking_mode/), and [vision](https://api-docs.deepseek.com/guides/vision/).

<script src="/public/feedback.js" />


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.