A Library for Comprehensive Research on Mixture of Experts in Large Language Models
MoE has become a core architecture behind frontier language and multimodal systems, but meaningful MoE research is still out of reach for many labs: large studies often require hundreds of H100/A100 GPUs, leaving the community with fragmented small-scale or synthetic comparisons. LibMoE lowers that barrier by packaging training, zero-shot evaluation, and routing analysis into a standardized pipeline that runs under constrained compute while preserving behaviors observed in larger MoE models.
Study modern MoE methods without frontier-lab training budgets.
Compare pretraining, sparse upcycling, and evaluation in one reproducible pipeline.
Analyze routing behavior that remains relevant to larger deployed MoE systems.
Framework
A modular, extensible infrastructure for systematic SMoE research — from small-scale pretraining to large-scale vision-language models — all within realistic compute budgets.
Unified abstraction for routing strategies with fully customizable gating, load balancing, and expert selection logic.
End-to-end pipelines for both LLM pretraining and VLM sparse upcycling with full parallelism support.
Plug-and-play evaluation across standardized benchmarks with routing instrumentation and expert analytics.
Empirical Results
Systematic comparison of 7 SMoE algorithms across language and vision-language tasks, all under resource-constrained settings accessible to the broader research community.
| Method | AI2D | TextVQA | GQA | MMBench | HalBench | MathVista | MMMU | MMStar | Pope | MME | MME RW | OCR | AVG Acc | AVG Rank |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SMoE | 69.56 | 43.93 | 61.51 | 71.31 | 46.90 | 37.90 | 41.56 | 41.23 | 86.28 | 63.33 | 27.83 | 37.50 | 52.40 | 5.96 |
| XMoE | 69.72 | 43.93 | 61.52 | 72.25 | 47.42 | 38.50 | 42.11 | 43.99 | 86.61 | 63.81 | 29.18 | 39.40 | 53.20 | 3.33 |
| σ-MoE | 69.79 | 44.69 | 61.70 | 71.74 | 47.00 | 38.40 | 43.11 | 42.08 | 86.69 | 63.80 | 29.70 | 38.40 | 53.09 | 3.33 |
| SharedE-V2 | 70.56 | 45.04 | 61.34 | 71.13 | 47.11 | 39.50 | 42.78 | 42.73 | 86.53 | 63.93 | 29.55 | 38.70 | 53.24 | 3.33 |
| SharedE-V3 | 71.92 | 44.59 | 61.93 | 72.59 | 46.37 | 38.20 | 42.33 | 43.30 | 86.78 | 64.61 | 29.29 | 39.40 | 53.44 | 2.46 |
| TC-MoE | 70.08 | 43.75 | 61.89 | 71.05 | 45.74 | 38.10 | 41.89 | 43.64 | 86.76 | 63.05 | 31.84 | 38.30 | 53.01 | 4.50 |
| MoE++ | 70.13 | 43.37 | 61.52 | 71.39 | 46.16 | 38.60 | 40.78 | 43.24 | 86.60 | 63.26 | 28.19 | 37.50 | 52.56 | 5.08 |
| Method | AI2D | TextVQA | GQA | MMBench | HalBench | MathVista | MMMU | MMStar | Pope | MME | MME RW | OCR | AVG Acc | AVG Rank |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SMoE | 65.52 | 41.51 | 61.62 | 72.25 | 41.75 | 29.50 | 41.67 | 42.24 | 87.13 | 61.05 | 32.15 | 31.70 | 50.67 | 3.42 |
| XMoE | 65.84 | 41.96 | 61.61 | 72.16 | 41.85 | 31.60 | 41.67 | 40.32 | 86.64 | 60.95 | 32.73 | 33.20 | 50.88 | 3.08 |
| σ-MoE | 65.52 | 42.08 | 61.74 | 71.22 | 40.80 | 30.30 | 41.22 | 41.99 | 86.64 | 61.44 | 32.31 | 32.50 | 50.65 | 3.71 |
| SharedE-V2 | 64.77 | 41.96 | 61.27 | 71.74 | 41.85 | 30.90 | 43.33 | 41.56 | 86.82 | 60.52 | 32.88 | 31.40 | 50.75 | 3.67 |
| SharedE-V3 | 65.58 | 42.06 | 61.26 | 72.42 | 41.43 | 30.60 | 42.44 | 41.75 | 86.81 | 60.93 | 31.47 | 32.60 | 50.86 | 3.33 |
| TC-MoE | 65.50 | 40.70 | 61.19 | 71.22 | 42.06 | 29.30 | 41.22 | 41.10 | 86.53 | 60.30 | 31.89 | 33.00 | 50.34 | 5.42 |
| MoE++ | 65.03 | 41.44 | 60.61 | 71.74 | 42.69 | 30.30 | 43.00 | 40.32 | 86.58 | 60.11 | 31.32 | 31.20 | 50.36 | 5.38 |
Note: MME represents sum of perception and cognition score. MME RW represents MME RealWorld. AVG Acc refers to average accuracy across all tasks. AVG Rank denotes the average ranking across all metrics.
| Method | PPL ↓ | LAMBADA | BLiMP | CBT | HellaSwag | PIQA | ARC-E | RACE | SIQA | CommonSenseQA | AVG Acc | AVG Rank |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SMoE | 13.63 | 25.27 | 77.71 | 84.18 | 29.43 | 57.94 | 32.68 | 30.11 | 35.62 | 24.65 | 44.18 | 4.50 |
| XMoE | 13.98 | 24.57 | 76.53 | 84.12 | 29.34 | 58.27 | 32.26 | 29.69 | 35.47 | 24.49 | 43.86 | 6.45 |
| σ-MoE | 13.61 | 25.43 | 77.38 | 84.23 | 29.13 | 58.92 | 32.73 | 31.05 | 34.90 | 24.90 | 44.30 | 4.30 |
| SharedE-V2 | 13.49 | 25.29 | 77.37 | 84.33 | 29.38 | 60.17 | 33.83 | 31.02 | 35.57 | 24.98 | 44.66 | 2.80 |
| SharedE-V3 | 13.42 | 25.49 | 77.20 | 84.40 | 29.38 | 59.14 | 32.52 | 30.60 | 35.57 | 25.47 | 44.42 | 3.20 |
| TC-MoE | 13.51 | 25.60 | 76.91 | 84.68 | 29.27 | 59.03 | 33.02 | 30.63 | 36.03 | 26.37 | 44.62 | 2.90 |
| MoE++ | 13.54 | 25.45 | 77.23 | 84.83 | 29.28 | 58.49 | 33.49 | 30.11 | 35.62 | 24.49 | 44.33 | 3.85 |
| Method | PPL ↓ | LAMBADA | BLiMP | CBT | HellaSwag | PIQA | ARC-E | RACE | SIQA | CommonSenseQA | AVG Acc | AVG Rank |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SMoE | 9.51 | 37.13 | 80.47 | 89.83 | 37.49 | 64.36 | 38.22 | 33.03 | 37.41 | 26.54 | 49.39 | 5.15 |
| XMoE | 9.66 | 35.25 | 80.38 | 89.35 | 37.19 | 64.20 | 38.99 | 32.95 | 37.77 | 28.34 | 49.38 | 5.75 |
| σ-MoE | 9.46 | 37.56 | 81.08 | 89.57 | 37.52 | 64.91 | 39.15 | 32.68 | 37.67 | 28.50 | 49.85 | 3.60 |
| SharedE-V2 | 9.52 | 37.11 | 80.98 | 89.93 | 37.14 | 64.36 | 38.06 | 33.17 | 36.95 | 27.35 | 49.45 | 5.25 |
| SharedE-V3 | 9.49 | 36.88 | 81.28 | 89.65 | 37.32 | 65.72 | 38.86 | 33.12 | 38.59 | 28.09 | 49.95 | 3.60 |
| TC-MoE | 9.38 | 37.87 | 81.21 | 90.19 | 37.95 | 64.47 | 39.28 | 33.77 | 37.92 | 27.85 | 50.06 | 2.35 |
| MoE++ | 9.38 | 38.80 | 80.88 | 89.77 | 37.70 | 64.64 | 39.37 | 34.02 | 37.97 | 28.34 | 50.16 | 2.30 |
Note: PPL represents perplexity (lower is better). LAMBADA, BLiMP, CBT, HellaSwag, PIQA, ARC-Easy, RACE, SIQA, and CommonSenseQA represent task accuracy (%). AVG Acc refers to average accuracy across all tasks. AVG Rank denotes the average ranking across all metrics.
Current SMoE methods achieve broadly similar performance under matched compute, and none delivers a clearly dominant accuracy gain relative to the extra routing complexity it introduces. The more practical signal is efficiency: newer designs such as SharedE and MoE++ remain competitive in AVG Acc while reducing training and inference cost, making them stronger choices when runtime and resource budgets matter.
Deep Analysis
LibMoE turns routing logs into regime-level diagnostics: whether experts keep changing, whether shared experts stabilize faster under sparse upcycling, and whether layer-wise confidence matches patterns seen in much larger VLMs.
The same probes are applied to from-scratch pretraining (0.15B LLM, 6.55B tokens), VLM sparse upcycling (LLaVA-665K), and Qwen3-VL-30B-A3B as a large-scale reference.
Router behavior stabilizes as training progresses: for the same input token, expert choices become less volatile across checkpoints. SharedE is less stable early in from-scratch pretraining because the router must align expert selection with shared experts that are also initialized from scratch and changing under the task loss; after this co-adaptation phase, routing stabilizes again.
Choosing the right top-K experts matters because it reveals whether a router is exploiting the capacity of its MoE architecture. When top experts are perturbed, vanilla SMoE can improve, suggesting unused routing capacity; newer methods generally drop, indicating that their routers already rely more tightly on the selected experts.
Routers do not need the same expert spread for every input. Reasoning-heavy tasks produce narrower expert allocation, indicating stronger specialization, while simpler or broader tasks keep routing more distributed. The useful behavior is adaptive specialization: concentrate when the task demands it, but preserve coverage elsewhere.
After experts are selected, the final output is usually not dominated by only one or two experts. Weight allocation remains relatively balanced across selected experts, meaning MoE decisions are often collaborative rather than a single-expert takeover.
Router initialization already shapes early load balance before any architectural change. A smaller initialization scale keeps logits less extreme, improves routing diversity, and reduces balance loss without adding new auxiliary-loss machinery.