Kimi and related services and features are operated by Moonshot AI PTE. LTD. and may include trademarks, logos, software, models, and other proprietary materials owned by Moonshot AI or its licensors. All rights reserved.
No One Cares About Your Bio: Change This Feature Today
A plain text description won't keep profile visitors engaged for more than two seconds. Natively feature your favorite song, album, or artist playlist using our integrated Spotify module.
MiniMax M3: The Open-Weight Multimodal Coding Model Built for 1M-Token Context
MiniMax M3 is an open-weight, natively multimodal model designed for coding, agentic workflows, and long-context reasoning, with support for up to 1 million tokens of context and a 428B-parameter Mixture-of-Experts architecture that activates about 23B parameters per token. It is positioned as a frontier model for developers who need sustained, multi-step task execution across code, tools, documents, images, and video.
What MiniMax M3 is
MiniMax describes M3 as a frontier coding and agentic model built on its MiniMax Sparse Attention architecture, or MSA, with a 1M-token context window and native multimodality. The model is intended for long-range coding, autonomous task decomposition, tool invocation, and long-video understanding.
In practical terms, that means M3 is not just a text model with added vision features. It is trained from the start to work across text, image, and video inputs, while keeping attention efficient enough to handle extremely large contexts.
Core technical characteristics
M3 combines several design choices that make it unusual among open-weight models.
- Architecture: 428B-parameter Mixture-of-Experts model with about 23B activated parameters per token.
- Context window: Up to 1M tokens, with MiniMax also stating a guaranteed minimum of 512K tokens for API use.
- Modality: Native support for text, image, and video input, with text output.
- Attention design: MiniMax Sparse Attention, created to reduce the cost of long-context inference.
The key technical story is that M3 aims to keep long-context work feasible without the full quadratic cost of standard attention. MiniMax and several platform listings state that MSA is central to how the model reaches 1M-token scale.
Why the 1M-token context matters
A 1M-token context window is valuable when a task involves large codebases, long logs, extended documents, or multi-hour agent workflows. MiniMax says the context capacity supports long-range coding, long-range agent tasks, and long-video understanding.
For developers, this can reduce the need to constantly chop up code, documentation, or history into small chunks. For agentic use cases, it also helps the model retain task state over many steps, which is important for tool use and autonomous execution.
Where MiniMax M3 stands out
MiniMax presents M3 as especially strong in coding and agentic benchmarks, and its own blog highlights performance on specialized tasks such as software engineering and tool-based workflows. Public model listings from NVIDIA and other platforms echo that positioning, describing M3 as a model for long-horizon coding, agentic work, and creative or design tasks.
- Coding: Built for long-horizon programming tasks and code understanding.
- Agent workflows: Designed for task decomposition, tool invocation, and multi-step reasoning.
- Multimodal input: Accepts image and video alongside text.
- Long-form media: Intended for long-video understanding, with some listings noting support for videos up to 30 minutes.
Reported benchmark performance
MiniMax’s blog reports the following benchmark results for M3: SWE-Bench Pro at 59.0%, Terminal-Bench 2.1 at 66.0%, and SWE-efficiency at 34.8%. These results are presented by the vendor as evidence of strong coding and agentic performance.
Because these are vendor-reported figures, they are useful for understanding MiniMax’s claims, but they should be interpreted alongside independent evaluations as they become available. Even so, the consistency of the model’s positioning across MiniMax, NVIDIA, and other infrastructure providers suggests that M3 is being treated as a serious open-weight contender in long-context coding.
What the model is designed to do
MiniMax and third-party model cards describe M3 as suitable for several demanding workflows.
- Long-horizon coding tasks that can span many hours
- Agentic work involving tool calls and iterative execution
- Search-heavy office workflows
- Long-form video understanding
- Design and creative workflows that combine text and visuals
Some listings also describe reasoning modes such as enabled and adaptive, indicating that the model can vary its level of deliberate reasoning depending on the task.
Why developers are paying attention
MiniMax M3 is notable because it combines three capabilities that are often hard to deliver together: open-weight availability, native multimodality, and ultra-long context. Many models do one or two of these well, but M3 is explicitly built around all three.
That makes it appealing for developers building coding assistants, document analyzers, multimodal agents, and long-running workflows that need memory across a very large span of inputs.
What to watch next
The most important questions for M3 over time will be how well it performs outside vendor benchmarks, how reliably it handles long autonomous tasks in production, and how efficiently it runs at scale in real deployments. Its architecture and positioning are promising, but broad adoption will depend on practical results in everyday developer workflows.
Conclusion
MiniMax M3 is best understood as a frontier open-weight model for developers who need long-context reasoning, coding, and multimodal understanding in one system. Its 1M-token context window, 428B MoE design, and native support for text, image, and video make it one of the most ambitious coding-focused models currently available.
The real shift is not the million-token window by itself, but the fact that M3 turns context into an execution substrate: if the model can preserve state across text, images, video, and tool calls, then the competitive advantage moves from prompting skill to workflow orchestration. In that sense, M3 is less a bigger model than an operating layer for agentic software.
What attention mechanism allows MiniMax M3 to handle 1M tokens?
