Your privacy choices

Allow optional cookies for referral attribution, visit analytics, and Google Ads purchase measurement.

Back to blog

MiniMax M3.1 Flash: 1M Context and API Guide

MiniMax M3.1MiniMax APIM3.1 FlashOpenAI-compatible APIAI coding
MiniMax M3.1 Flash: 1M Context and API Guide cover

MiniMax's newest M-series route is MiniMax-M3.1-flash. It targets agentic reasoning, tool use, coding, and long-context work. It supports up to 1,000,000 tokens of context, text/image/video input, and tunable thinking depth.

The model ID matters as much as the model name. This site's current catalog uses the new MiniMax-M3.1-flash route; do not copy the Preview suffix from older materials into production requests. Check account access, the live model list, and current pricing before changing production code.

This guide follows the MiniMax official model invocation documentation. Vendor claims, independent benchmarks, and this site's gateway availability are separate questions. For this site, model IDs, routing, and billing are shown on the live model pricing page.

MiniMax M3.1 Flash specifications

Item Official information
Exact model ID MiniMax-M3.1-flash
Context window 1,000,000 tokens
Input Text, images, and video
Output Text
Thinking Always enabled; cannot be disabled
Thinking depth low, medium, high, xhigh, max
Default effort max
Recommended protocol Anthropic-compatible
OpenAI-compatible base URL https://api.minimax.io/v1
Anthropic-compatible base URL https://api.minimax.io/anthropic
Current site route MiniMax-M3.1-flash is listed in the live catalog

A 1M-token context is useful for long documents, full repositories, and multi-step agent sessions. It does not mean every request should use the maximum window. Longer inputs can increase latency, output requirements, and actual usage, so measure the tasks that matter to you.

What changes from MiniMax M3?

Both models target long context, multimodal input, and coding workflows. Migration still requires more than changing the model string:

Capability M3.1 Flash MiniMax M3
Thinking Always enabled Can be enabled or disabled
Thinking depth low through max No effort levels described here
OpenAI-compatible field reasoning_effort; thinking is in reasoning_content Follow the current response shape
Context 1M 1M
Input Text, image, and video Official page lists image and video input
Prompt caching Supported Supported

If older code sends thinking: {"type": "disabled"} or effort: "none", M3.1 returns a 400 error. Lower the effort level when you need less latency or fewer thinking tokens; do not try to turn thinking off.

Calling MiniMax M3.1 with the OpenAI SDK

MiniMax documents an OpenAI-compatible path. Install the SDK and keep the key in an environment variable:

pip install openai
export MINIMAX_API_KEY="your_api_key"

Minimal Chat Completions example:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.minimax.io/v1",
    api_key=os.environ["MINIMAX_API_KEY"],
)

response = client.chat.completions.create(
    model="MiniMax-M3.1-flash",
    reasoning_effort="high",
    messages=[
        {"role": "user", "content": "Explain an API retry boundary in three sentences."},
    ],
)

print(response.choices[0].message.content)

reasoning_effort accepts low, medium, high, xhigh, and max. In the OpenAI-compatible response, thinking is returned in reasoning_content and the final answer is in content. Application code should tolerate a missing reasoning_content field and only store or display it when needed.

MiniMax also lists an Anthropic-compatible SDK as the recommended path:

import anthropic

client = anthropic.Anthropic(
    base_url="https://api.minimax.io/anthropic",
    api_key="your_api_key",
)

message = client.messages.create(
    model="MiniMax-M3.1-flash",
    output_config={"effort": "high"},
    max_tokens=4096,
    messages=[{"role": "user", "content": "Check the edge cases in this code."}],
)

The field names differ by protocol: Anthropic-compatible requests use output_config.effort, while OpenAI-compatible requests use reasoning_effort. Check thinking blocks, output limits, and streaming events together; changing only the base URL is not a complete migration.

Workloads to test first

  • Repository maintenance: provide the repository, reproduction steps, acceptance criteria, and test command, then check whether the model completes the diagnose, edit, and verify loop.
  • Long-document research: provide sources, dates, and citation rules, then review how it separates evidence, assumptions, and conclusions.
  • Multimodal analysis: attach an image or video and request a structured inspection. Image or video input does not mean native image or video output.
  • Tool use: start with a reversible task and verify arguments, errors, and repeated calls before expanding permissions or context.

Official examples run in specific environments and are not guarantees for every project. Record the model ID, effort level, input/output tokens, elapsed time, and retry count before promoting a workflow.

Migration checklist from M3

  1. Replace the old model name with MiniMax-M3.1-flash and confirm account access.
  2. Remove parameters that disable thinking; set reasoning_effort or output_config.effort explicitly.
  3. Update response parsing to tolerate missing reasoning_content and read final text from content.
  4. Reserve enough output budget for both reasoning and the final answer.
  5. Test text, image, video, tool calls, streaming, and cache behavior separately.
  6. Compare completion rate, total latency, token usage, and the final bill on the same task set.

FAQ

Why do older materials still show a Preview suffix?

Some provider materials still use the Preview suffix, but this site's current public catalog uses the new MiniMax-M3.1-flash route ID. Actual access and billing still depend on account permissions, the live catalog, and the bill.

Can I disable thinking on M3.1 Flash?

No. MiniMax says thinking is always enabled, and requests that disable it or use effort: "none" return 400. Choose a lower effort level when latency matters.

How large is the MiniMax M3.1 context window?

The official model table lists 1,000,000 tokens. Usable length is still affected by the endpoint, output budget, client limits, and account policy.

Does MiniMax M3.1 support the OpenAI SDK?

Yes. The official guide shows https://api.minimax.io/v1 with Chat Completions. OpenAI compatibility does not guarantee every OpenAI feature, so validate tools, streaming, and Responses-style requests separately.

Can I call M3.1 through this site?

The public catalog changes with provider routing and account access. Check the live model pricing page and run a small request with the exact displayed model ID. An official release does not by itself prove that a third-party gateway has synchronized the route.

Sources

Sources checked on October 1, 2026. This is an integration and selection guide, not an independent benchmark.