一句话摘要:这篇论文给出了”MoE 的能力落在哪两个稠密模型之间”的定量答案:DeepSeekMoE 16B ≈ LLaMA2 7B 的质量但只用约 40% 算力;145B ≈ DeepSeek 67B 稠密但只用 28.5% 算力;而同总参数量的稠密模型是 MoE 的能力上界。
原始标题
DeepSeekMoE: Towards Ultimate Expert Specialization in MoE Language Models(DeepSeek-AI,arXiv:2401.06066,2024-01-11)
来源:https://arxiv.org/abs/2401.06066
它解决什么问题
传统 MoE(如 GShard)从 N 个专家里选 top-K 个,但存在两个毛病:专家专业化程度不足(多个专家学到重叠知识),以及存在冗余(通用知识被反复学)。DeepSeekMoE 提出两条策略来解决。
关键要点
1. 两条策略
“(1) finely segmenting the experts into mN ones and activating mK from them, allowing for a more flexible combination of activated experts; (2) isolating K_s experts as shared ones, aiming at capturing common knowledge and mitigating redundancy in routed experts.”
- 细粒度专家分割(fine-grained expert segmentation):把专家切得更碎、激活更多个,让激活组合数爆炸式增长。
- 共享专家隔离(shared expert isolation):留几个专家对每个 token 都激活,专门承载通用知识;其余路由专家只需学专门化的部分。
2. 三个可以直接引用的定量结论
“Starting from a modest scale with 2B parameters, we demonstrate that DeepSeekMoE 2B achieves comparable performance with GShard 2.9B, which has 1.5 times the expert parameters and computation. In addition, DeepSeekMoE 2B nearly approaches the performance of its dense counterpart with the same number of total parameters, which set the upper bound of MoE models. Subsequently, we scale up DeepSeekMoE to 16B parameters and show that it achieves comparable performance with LLaMA2 7B, with only about 40% of computations. Further, our preliminary efforts to scale up DeepSeekMoE to 145B parameters consistently validate its substantial advantages over the GShard architecture, and show its performance comparable with DeepSeek 67B, using only 28.5% (maybe even 18.2%) of computations.”
从 20 亿参数这个不算大的规模开始,DeepSeekMoE 2B 就能达到 GShard 2.9B 的水平(后者用了 1.5 倍的专家参数和计算量);而且它几乎追平了同总参数的稠密模型——论文指出,这个稠密模型设定了 MoE 的能力上界。
| 对比 | 结论 |
|---|---|
| DeepSeekMoE 2B vs 同总参数稠密模型 | 几乎追平 → 同总参数稠密模型 = MoE 的能力上界 |
| DeepSeekMoE 16B vs LLaMA2 7B(稠密) | 质量相当,只用约 40% 算力 |
| DeepSeekMoE 145B vs DeepSeek 67B(稠密) | 质量相当,只用 28.5% 算力 |
3. 这三条合起来回答了什么
MoE 的能力区间:高于”同激活参数量的稠密模型”,低于”同总参数量的稠密模型”。所以”37B 激活 = 37B 稠密”这个说法不成立,但”37B 激活 > 37B 稠密”也不成立——正确说法是夹在中间,且随算力预算增大,MoE 相对稠密的效率优势扩大。
4. 补充:MoE 缩放定律的正反方
正方(Krajewski et al., arXiv:2402.07871):
“Our results suggest that a compute-optimal MoE model trained with a budget of 10^20 FLOPs will achieve the same quality as a dense Transformer trained with a 20× greater computing budget, with the compute savings rising steadily, exceeding 40× when budget of 10^25 FLOPs is surpassed.”
我们的结果显示:一个用 10^20 FLOPs 算力训练的计算最优 MoE 模型,能达到「用 20 倍算力训练的稠密 Transformer」同等的质量;而且这个节省幅度会持续扩大,预算超过 10^25 FLOPs 时超过 40 倍。
反方(同一篇论文引用的对立结论):
“other studies have stated that the gap in efficiency between MoE and standard Transformers narrows at scale (Artetxe et al., 2022) or even that traditional dense models may outperform MoE as the size of the models increases (Clark et al., 2022).”
也有研究认为,MoE 与标准 Transformer 的效率差距会随规模缩小,甚至在大到一定程度后稠密模型可能反超。这是 MoE 路线的反方证据,值得一起看。
重要引用(英文原文)
“However, conventional MoE architectures like GShard, which activate the top-K out of N experts, face challenges in ensuring expert specialization, i.e. each expert acquires non-overlapping and focused knowledge.”
但 GShard 这类传统 MoE 架构(从 N 个专家里选 top-K)在确保专家专业化上存在困难——也就是让每个专家都学到互不重叠、足够聚焦的知识。
边界与争议
- 145B 这一档是 “preliminary efforts”(初步尝试),论文自己也标注了不确定性(“28.5% (maybe even 18.2%)”)。
- “同总参数稠密模型是上界” 这个结论是基于 2B 规模的实验得出的,论文没有在 145B 规模上重做这个对照。
- 反方文献(Artetxe 2022、Clark 2022)确实存在,正方论文的解释是”那些结论依赖了固定训练时长和固定专家大小的隐含假设”。
与本文其他页面的关系
- 混合专家 —— 概念页,讲机制
- 激活参数 —— 概念页,讲”激活参数不等于能力”
- DeepSeek-V3技术报告 —— 把这套架构推到 671B 的工程实现