Back to papers
March 19, 2026cs.LGAdvanced
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
AI-Generated Summary
This paper presents DyMoE, a technique to make large AI models with multiple expert components run efficiently on edge devices (like mobile phones or IoT devices) by intelligently compressing less important experts while keeping critical ones intact. The method uses dynamic compression strategies that adapt based on which experts matter most and where in the model they're located, achieving 3-22x faster inference speeds compared to existing approaches while maintaining accuracy.
Difficulty
Advanced
Categories
cs.LG
AI Tags
model compressionquantizationmixture of expertsedge inferenceoptimizationefficient AI