MoE-CAP
收藏资源简介:
MoE-CAP是一个专门为评估MoE系统而设计的基准,旨在理解和评估MoE系统的成本、准确性和性能。该数据集包含多种MoE模型,并使用新的稀疏感知性能指标进行评估,包括稀疏内存带宽利用率和稀疏模型FLOPS利用率。MoE-CAP还引入了CAP雷达图,以直观地展示MoE系统在成本、准确性和性能方面的权衡。
MoE-CAP is a benchmark specifically tailored for evaluating Mixture-of-Experts (MoE) systems, aiming to comprehensively understand and assess the cost, accuracy, and performance of such systems. This dataset encompasses a variety of MoE models, and employs novel sparse-aware performance metrics for evaluation, including sparse memory bandwidth utilization and sparse model FLOPS utilization. Additionally, MoE-CAP introduces the CAP Radar Chart to intuitively demonstrate the trade-offs of MoE systems across cost, accuracy, and performance dimensions.
数据集概述:MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
基本信息
- 标题: MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
- arXiv标识符: arXiv:2505.11415v1
- 提交日期: 2025年5月16日
- 作者: Yinsicheng Jiang, Yao Fu, Yeqi Huang, Ping Nie, Zhan Lu, Leyang Xue, Congjie He, Man-Kit Sit, Jilong Xue, Li Dong, Ziming Miao, Dayou Du, Tairan Xu, Kai Zou, Edoardo Ponti, Luo Mai
- DOI: https://doi.org/10.48550/arXiv.2505.11415
摘要
稀疏混合专家(MoE)架构在高效扩展大型语言模型(LLMs)方面越来越受青睐,但其依赖于异构计算和内存资源。这些因素共同影响系统的成本(Cost)、准确性(Accuracy)和性能(Performance)(CAP),使得权衡不可避免。现有的基准测试往往无法准确捕捉这些权衡,使实际部署决策复杂化。为此,我们引入了MoE-CAP,一个专为MoE系统设计的基准测试。我们的分析表明,在当前硬件上实现CAP的平衡是困难的;MoE系统通常优化三个维度中的两个,而牺牲第三个——我们称之为MoE-CAP权衡。为了可视化这一点,我们提出了CAP雷达图。我们还引入了稀疏感知性能指标——稀疏内存带宽利用率(S-MBU)和稀疏模型FLOPS利用率(S-MFU)——以支持在不同硬件平台和部署场景下对MoE系统进行准确的性能基准测试。
学科分类
- 主要学科: 机器学习(cs.LG)
- 次要学科: 分布式、并行和集群计算(cs.DC)
相关链接
- PDF链接: http://arxiv.org/pdf/2505.11415v1
- HTML链接: http://arxiv.org/html/2505.11415v1
- TeX源码: http://arxiv.org/format/2505.11415v1
提交历史
- 版本1: 2025年5月16日提交,文件大小284 KB

- 1MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems爱丁堡大学, 微软研究院, 腾讯, NetMind.AI, 英伟达 · 2025年



