MiniMax-M2 系列:小激活释放最大现实智能

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

闫俊杰 Junjie Yan · MiniMax · 2026-05-26 · arXiv:2605.26494 ↗ · 被引 9

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 MiniMax-M2 系列,这是一个基于混合专家模型的语言模型家族,其核心理念是小激活可以释放最大的现实智能。旗舰模型 M2 总参数量为 2299 亿,但每个 token 仅激活 98 亿参数。M2 系列专为智能体部署而设计,基于三个组件:(i) 智能体驱动的数据管道,在智能体编码和智能体协作中生成大规模、可验证的轨迹,每个轨迹都基于可执行工作空间和与工件对齐的奖励;(ii) Forge,一个可扩展的智能体原生强化学习系统,适应长程智能体轨迹,并配备窗口式 FIFO 调度、前缀树合并、推理优化,以及支持白盒和黑盒智能体的干净训练-推理-智能体解耦;(iii) 最新的 M2.7 检查点向自我进化迈出了早期一步——自主调试训练运行并修改自身框架。从 M2 到 M2.7,这种组合将小激活足迹转化为在智能体编码、深度搜索、办公任务和推理基准上的前沿性能。

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 21)

阅读逐段中英对照全文 →