On June 1, 2026, the domestic AI unicorn MiniMax officially released a new generation of flagship large model MiniMax M3. This model is officially positioned as the first domestic open source flagship with three core capabilities: "cutting-edge programming and agent capabilities, million-level contexts, and native multimodal". It is directly benchmarked with GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro and other overseas closed-source top models.
1. Three core competencies: the "three-in-one" flagship that breaks overseas monopoly
1. Cutting-edge programming and agent capabilities: code is directly deliverable
The MiniMax M3 has reached the industry-leading level in coding capabilities. In the SWE-Bench Pro evaluation, which measures real software engineering capabilities, the M3 score surpassed GPT-5.5 and Gemini 3.1 Pro, close to Claude Opus 4.7; in Terminal Bench 2.1, it led Opus 4.7 (64.1) and GPT-5.5 (58.6) with 66.0 points.

Officials stressed thatThe code goal written by M3 is "directly deliverable" rather than "can run but need to be changed." In the BrowseComp agent evaluation, M3 surpassed Opus 4.7 (79.3) with 83.5 points, demonstrating strong autonomous browsing and information retrieval capabilities.
A very convincing measured example is:MiniMax handed over an outstanding ICLR 2025 paper to M3 for independent reproduction. M3 ran continuously for nearly 12 hours, independently producing 18 commits and 23 experimental charts throughout the entire process, successfully running through the core experiment. Multimodal capabilities allow it to understand charts and formulas in papers, millions of contexts ensure that papers + code + experimental logs enter the window at one time, and programming +Agent capabilities drive long thread execution.

2. Million level context: 1M tokens '"long-range infrastructure"
Based on self-researchWith the MiniMax Sparse Attention (MSA) architecture, the M3 API supports up to 1M tokens context windows and ensures that at least 512K tokens are available. The official made it clear that the 1M context is the "infrastructure" for long-term Agent, long-term Coding, and long-term video understanding.
The MSA architecture combines Index Branch fast indexing and Sparse Branch precise computing to solve the bottleneck of square growth in computing complexity when traditional transformers process millions of Tokens. Compared with the previous generation M2, the M3 achieves 9.7x acceleration in the Prefill stage and 15.6x acceleration in the Decoding stage; in the 1M context, the calculation amount per Token is only 1/20 of that of the previous generation model.
This efficiency leap means that companies process millionsThe computing power cost of long Token-level documents can be reduced by more than 80%, and individual users respond with almost no delay when engaging in long conversations.
3. Native multimodal: Visual alignment starting from step zero
M3 is a native multimodal model that reconstructs the entire data pipeline, expands the pre-training data scale to the order of 100 T, and starts multimodal training from step 0 to make the text and visual semantic space highly aligned. On the OmniDocBench multimodal test set, M3 scored better than the Gemini 3.1 Pro; in the SVG-Bench comprehensive evaluation, M3 surpassed Opus 4.7 (62.3) with 63.7 points.
2. Architectural innovation: MSA-"minimalism" that leads engineering
M3's core architectural innovation is MiniMax Sparse Attention (MSA), a new sparse attention architecture that is simple and easy to scale.
Compared with traditional dense attention mechanisms,The design of MSA reflects the rational trade-off of "engineering first":
Underlayer of GQA rather than MLA: Using Group Query Attention (GQA) means that the kernel functions of vLLM, SGLang, and FlashAttention can be reused with almost zero modification, without the need for complex engineering modifications for implicit KVs.
Block-level selection, real KV calculation: Unlike the solution of calculating attention on compressed KV, M3 retains the full expression ability of Softmax's attention, exchanging Token economics for model quality.
Single-branch minimalist design: Compared with DeepSeek NSA's three-way parallelism (compression + selection + sliding window)+ learning gating, M3 only retains select branches, cuts off redundant components, and pursues "immediate and fast running" pragmatism.
At the operator level,MSA uses KV blocks as the outer layer to aggregate KV outer gather Q that hit the query. Each block is read only once and has continuous memory access. The calculation memory access ratio is significantly better than the common method, and more than 4 times faster than the open source Flash-Sparse-Attention and flash-moba.
3. Limit measurement: With no intervention for 24 hours, FP8 utilization rate soared from 7.6% to 71.3%
M3 demonstrates strong autonomous planning capabilities for long threads in extreme tasks. In addition to reproducing ICLR papers for 12 hours, M3 also ran continuously for 24 hours without reference code and called tools nearly 2000 times, increasing the FP8 matrix multiplication hardware utilization on the Hopper architecture from 7.6% to 71.3%.
In In the PostTrainBench open evaluation, M3 was given four Base models that only completed pre-training, and was required to independently complete the entire process of data synthesis, training, evaluation, and iteration within 12 hours. With no intervention throughout the process, M3 finally scored 37.1, ranking third, behind Opus 4.7 (42.4) and GPT-5.5 (39.3).
4. Open source and pricing: breaking the "high-price monopoly" cost-effective revolution
Will soon be fully open source
MiniMax promises that M3 will soon be open source on HuggingFace and GitHub, supporting private cluster deployment and fine-tuning. This will make it the first open source model in China that has the three elements of "cutting-edge programming + million contexts + native multimodal", breaking the cutting-edge capability pattern previously monopolized by overseas closed-source models.
Competitive API pricing
M3 provides two versions of the API, standard version and M3-highspeed version, with exactly the same results, the latter being faster. Fully supports automatic caching, which takes effect automatically without setting.
Context ≤ 512K limited time limit 50% discount price:
Input: 2.1 yuan/million tokens for standard version, 3.15 yuan for priority version
Output: 8.4 yuan/million tokens for standard version, 12.6 yuan for priority version
Cache reading: 0.42 yuan/million tokens for standard version, 0.63 yuan for priority version
This price level is in sharp contrast to overseas flagships. As a reference,The input price of Claude Opus 4.6 is approximately US$15/million tokens (approximately RMB 108), and the output is as high as US$75 (approximately RMB 540). MiniMax's pricing strategy continues its "cost-effective" competitive route-according to Goldman Sachs data, MiniMax's current text API gross profit margin is about 40%, which is significantly higher than the industry average of 20%-25%.
Token Plan subscription upgrade
Synchronously upgradedToken Plan adopts a points-based usage deduction mechanism, covering three levels: Plus (¥49/month), Max (¥119/month) and Ultra (¥469/month). Core changes include conversion of points based on actual resource consumption, unified quota pool, and visual display of usage progress bars. Old plan users will also receive one-time compensation points.
5. Industry significance: The turning point from "parameter competition" to "efficiency competition"
The release of the MiniMax M3 marks a new stage in the competition for domestic large models:
Technically, M3 demonstrates that cutting-edge capabilities can be achieved through architectural innovation rather than simply heap parameters. Its MSA architecture, together with DeepSeek's NSA and Xiaomi's HySparse, form a domestic sparse attention technology matrix, pushing the industry from a "parameter scale competition" to a "efficiency and practicality competition."
On the commercial level, the release of the M3 coincides with the accelerated commercialization of MiniMax. The company's ARR has doubled since February, and the proportion of open platform/API revenue has increased from 30% at the end of 2025 to more than 50% currently. Institutions such as Morgan Stanley have given an "overallocation" rating and believe that the M3 series will bring "major intelligent breakthroughs."
On the ecological level, M3's open source commitment will greatly reduce the threshold for developers to use cutting-edge multimodal and long-context capabilities. In the 2026 technology map, the 1M context is changing from a "selling point" to a "baseline", while MiniMax is reshaping the cost-effective boundary of the domestic developer ecosystem in the form of all-factor open source.
Conclusion
The debut of the MiniMax M3 is not only an iteration of a model, but also a declaration of domestic AI in the era of "intelligent density". While overseas closed-source models are still monopolizing cutting-edge capabilities at high prices, M3's open source posture, new architecture, and excellent prices prove that domestic large models already have the ability to define industry standards. As weights on HuggingFace and GitHub are about to be opened, this "efficiency revolution" driven by the MSA architecture may profoundly change the tool choices of AI developers around the world.