TMLR 2026 Open Source 4 × H100

LibMoE

A Library for Comprehensive Research on Mixture of Experts in Large Language Models

Nam V. Nguyen · Thong T. Doan · Luong Tran · Van Nguyen · Quang Pham

MoE has become a core architecture behind frontier language and multimodal systems, but meaningful MoE research is still out of reach for many labs: large studies often require hundreds of H100/A100 GPUs, leaving the community with fragmented small-scale or synthetic comparisons. LibMoE lowers that barrier by packaging training, zero-shot evaluation, and routing analysis into a standardized pipeline that runs under constrained compute while preserving behaviors observed in larger MoE models.

01

Study modern MoE methods without frontier-lab training budgets.

02

Compare pretraining, sparse upcycling, and evaluation in one reproducible pipeline.

03

Analyze routing behavior that remains relevant to larger deployed MoE systems.

7
Algorithms
3
Model Scales
4×
H100 GPUs
44h
Max Training
Sparse Mixture of Experts Input Token x Router R(x, W_r) TopK · softmax · sparsity affinity scores s_R ∈ ℝᴺ Expert 1 g(x,W_e1) Expert 2 g(x,W_e2) Expert 3 Expert 4 Expert 5 g(x,W_e5) K=2 active · N=5 experts · sparsity = 60% ŷ = Σ s_R^i · g(x; W_ei) MoE Output ŷ weighted expert combination LibMoE supports 7 routing variants

Designing LibMoE

A modular, extensible infrastructure for systematic SMoE research — from small-scale pretraining to large-scale vision-language models — all within realistic compute budgets.

🔀

MoE Module

Unified abstraction for routing strategies with fully customizable gating, load balancing, and expert selection logic.

  • TopK sparse routing with custom K
  • Gating: softmax, sigmoid (σ-MoE)
  • 7 SOTA routing algorithms
  • Load balance loss variants
🏋️

Training Module

End-to-end pipelines for both LLM pretraining and VLM sparse upcycling with full parallelism support.

  • LLM: 0.15B → 0.68B (Switch-style)
  • VLM: 5.67B sparse upcycling
  • Data, Tensor, Pipeline, Expert parallelism
  • Runs in 6–44h on 4×H100
📊

Evaluation Module

Plug-and-play evaluation across standardized benchmarks with routing instrumentation and expert analytics.

  • LM: HellaSwag, ARC, WinoGrande…
  • VLM: MMStar, MMMU, MathVista…
  • Routing entropy & change rate tools
  • Expert similarity heatmaps
Supported Algorithms
SMoE σ-MoE TC-MoE XMoE MoE++ SharedE-V2 SharedE-V3

Benchmark Performance

Systematic comparison of 7 SMoE algorithms across language and vision-language tasks, all under resource-constrained settings accessible to the broader research community.

5.67B
Parameters (Sparse Upcycling)
53.44%
Best Avg Acc (SharedE-V3 1.2M)
50.88%
Best Avg Acc (XMoE 665K)
35h
MoE Training Time
LLaVA + OneVision (1.2M Samples) ViT Backbone · 5.67B Parameters · K=3 Active Experts
Method AI2D TextVQA GQA MMBench HalBench MathVista MMMU MMStar Pope MME MME RW OCR AVG Acc AVG Rank
SMoE 69.56 43.93 61.51 71.31 46.90 37.90 41.56 41.23 86.28 63.33 27.83 37.50 52.40 5.96
XMoE 69.72 43.93 61.52 72.25 47.42 38.50 42.11 43.99 86.61 63.81 29.18 39.40 53.20 3.33
σ-MoE 69.79 44.69 61.70 71.74 47.00 38.40 43.11 42.08 86.69 63.80 29.70 38.40 53.09 3.33
SharedE-V2 70.56 45.04 61.34 71.13 47.11 39.50 42.78 42.73 86.53 63.93 29.55 38.70 53.24 3.33
SharedE-V3 71.92 44.59 61.93 72.59 46.37 38.20 42.33 43.30 86.78 64.61 29.29 39.40 53.44 2.46
TC-MoE 70.08 43.75 61.89 71.05 45.74 38.10 41.89 43.64 86.76 63.05 31.84 38.30 53.01 4.50
MoE++ 70.13 43.37 61.52 71.39 46.16 38.60 40.78 43.24 86.60 63.26 28.19 37.50 52.56 5.08
LLaVA (665K Samples) ViT Backbone · 5.67B Parameters · K=3 Active Experts
Method AI2D TextVQA GQA MMBench HalBench MathVista MMMU MMStar Pope MME MME RW OCR AVG Acc AVG Rank
SMoE 65.52 41.51 61.62 72.25 41.75 29.50 41.67 42.24 87.13 61.05 32.15 31.70 50.67 3.42
XMoE 65.84 41.96 61.61 72.16 41.85 31.60 41.67 40.32 86.64 60.95 32.73 33.20 50.88 3.08
σ-MoE 65.52 42.08 61.74 71.22 40.80 30.30 41.22 41.99 86.64 61.44 32.31 32.50 50.65 3.71
SharedE-V2 64.77 41.96 61.27 71.74 41.85 30.90 43.33 41.56 86.82 60.52 32.88 31.40 50.75 3.67
SharedE-V3 65.58 42.06 61.26 72.42 41.43 30.60 42.44 41.75 86.81 60.93 31.47 32.60 50.86 3.33
TC-MoE 65.50 40.70 61.19 71.22 42.06 29.30 41.22 41.10 86.53 60.30 31.89 33.00 50.34 5.42
MoE++ 65.03 41.44 60.61 71.74 42.69 30.30 43.00 40.32 86.58 60.11 31.32 31.20 50.36 5.38

Note: MME represents sum of perception and cognition score. MME RW represents MME RealWorld. AVG Acc refers to average accuracy across all tasks. AVG Rank denotes the average ranking across all metrics.

66 / 8
Total / Active Experts
44.66%
Best Avg Acc (SharedE-V2 0.15B)
50.16%
Best Avg Acc (MoE++ 0.68B)
6–43h
LM Training Time
Small Model (0.15B Parameters) Pretraining Setting · 66 Total Experts · K=8 Active Experts
Method PPL ↓ LAMBADA BLiMP CBT HellaSwag PIQA ARC-E RACE SIQA CommonSenseQA AVG Acc AVG Rank
SMoE 13.63 25.27 77.71 84.18 29.43 57.94 32.68 30.11 35.62 24.65 44.18 4.50
XMoE 13.98 24.57 76.53 84.12 29.34 58.27 32.26 29.69 35.47 24.49 43.86 6.45
σ-MoE 13.61 25.43 77.38 84.23 29.13 58.92 32.73 31.05 34.90 24.90 44.30 4.30
SharedE-V2 13.49 25.29 77.37 84.33 29.38 60.17 33.83 31.02 35.57 24.98 44.66 2.80
SharedE-V3 13.42 25.49 77.20 84.40 29.38 59.14 32.52 30.60 35.57 25.47 44.42 3.20
TC-MoE 13.51 25.60 76.91 84.68 29.27 59.03 33.02 30.63 36.03 26.37 44.62 2.90
MoE++ 13.54 25.45 77.23 84.83 29.28 58.49 33.49 30.11 35.62 24.49 44.33 3.85
Large Model (0.68B Parameters) Pretraining Setting · 66 Total Experts · K=8 Active Experts
Method PPL ↓ LAMBADA BLiMP CBT HellaSwag PIQA ARC-E RACE SIQA CommonSenseQA AVG Acc AVG Rank
SMoE 9.51 37.13 80.47 89.83 37.49 64.36 38.22 33.03 37.41 26.54 49.39 5.15
XMoE 9.66 35.25 80.38 89.35 37.19 64.20 38.99 32.95 37.77 28.34 49.38 5.75
σ-MoE 9.46 37.56 81.08 89.57 37.52 64.91 39.15 32.68 37.67 28.50 49.85 3.60
SharedE-V2 9.52 37.11 80.98 89.93 37.14 64.36 38.06 33.17 36.95 27.35 49.45 5.25
SharedE-V3 9.49 36.88 81.28 89.65 37.32 65.72 38.86 33.12 38.59 28.09 49.95 3.60
TC-MoE 9.38 37.87 81.21 90.19 37.95 64.47 39.28 33.77 37.92 27.85 50.06 2.35
MoE++ 9.38 38.80 80.88 89.77 37.70 64.64 39.37 34.02 37.97 28.34 50.16 2.30

Note: PPL represents perplexity (lower is better). LAMBADA, BLiMP, CBT, HellaSwag, PIQA, ARC-Easy, RACE, SIQA, and CommonSenseQA represent task accuracy (%). AVG Acc refers to average accuracy across all tasks. AVG Rank denotes the average ranking across all metrics.

Benchmark Takeaway

Baseline Convergence & Complexity Gaps

Current SMoE methods achieve broadly similar performance under matched compute, and none delivers a clearly dominant accuracy gain relative to the extra routing complexity it introduces. The more practical signal is efficiency: newer designs such as SharedE and MoE++ remain competitive in AVG Acc while reducing training and inference cost, making them stronger choices when runtime and resource budgets matter.

Routing Dynamics
Under the Microscope

LibMoE turns routing logs into regime-level diagnostics: whether experts keep changing, whether shared experts stabilize faster under sparse upcycling, and whether layer-wise confidence matches patterns seen in much larger VLMs.

The same probes are applied to from-scratch pretraining (0.15B LLM, 6.55B tokens), VLM sparse upcycling (LLaVA-665K), and Qwen3-VL-30B-A3B as a large-scale reference.

01

Regime Shapes Stability

Router behavior stabilizes as training progresses: for the same input token, expert choices become less volatile across checkpoints. SharedE is less stable early in from-scratch pretraining because the router must align expert selection with shared experts that are also initialized from scratch and changing under the task loss; after this co-adaptation phase, routing stabilizes again.

02

Expert Selection Must Use Capacity

Choosing the right top-K experts matters because it reveals whether a router is exploiting the capacity of its MoE architecture. When top experts are perturbed, vanilla SMoE can improve, suggesting unused routing capacity; newer methods generally drop, indicating that their routers already rely more tightly on the selected experts.

03

Specialization Is Task-Dependent

Routers do not need the same expert spread for every input. Reasoning-heavy tasks produce narrower expert allocation, indicating stronger specialization, while simpler or broader tasks keep routing more distributed. The useful behavior is adaptive specialization: concentrate when the task demands it, but preserve coverage elsewhere.

04

Expert Weights Stay Collaborative

After experts are selected, the final output is usually not dominated by only one or two experts. Weight allocation remains relatively balanced across selected experts, meaning MoE decisions are often collaborative rather than a single-expert takeover.

05

Initialization Is a Cheap Balancing Lever

Router initialization already shapes early load balance before any architectural change. A smaller initialization scale keeps logits less extreme, improves routing diversity, and reduces balance loss without adding new auxiliary-loss machinery.

Research Team

Nam V. Nguyen
FPT Software AI Center
Thong T. Doan
FPT Software AI Center
Luong Tran
FPT Software AI Center
Van Nguyen
FPT Software AI Center
Quang Pham
Independent Researcher
📌 BibTeX Citation
@inproceedings{Nguyen2024LIBMoEAL,
  title={LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models},
  author={Nam V. Nguyen and Thong T. Doan and Luong Tran and Van Nguyen and Quang Pham},
  year={2024},
  url={https://api.semanticscholar.org/CorpusID:273812436}
}