Sitemap

Review: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (MoE)

6 min readDec 18, 2021

--

@ Medium)

Press enter or click to view image in full size
Sparsely-Gated Mixture-of-Experts Layer (MoE)
Press enter or click to view image in full size

The goal to train a trillion-parameter model on a trillion-word corpus.

Press enter or click to view image in full size
Model comparison on 1-Billion-Word Language-Modeling Benchmark
Press enter or click to view image in full size
Summary of high-capacity MoE-augmented models with varying computational budgets, vs. best previously published results
Language modeling on a 100 billion word corpus
Results on WMT’14 En>Fr newstest2014
Results on WMT’14 En>De newstest2014

The proposed approach achieved BLEU scores of 40.56 and 26.03 on the WMT’14 En>Fr and En>De benchmarks, outperforms GNMT and Deep-Att.

Press enter or click to view image in full size
Results on the Google Production En>Fr dataset

On the Google Production dataset, MoE model achieved 1.01 higher test BLEU score even after training for only one sixth of the time.

Multilingual Machine Translation

The MoE model achieves 19% lower perplexity on the dev set than the multilingual GNMT model.

--

--

Sik-Ho Tsang
Sik-Ho Tsang

PhD, Researcher. I share what I learn. :) Linktree: https://linktr.ee/shtsang for Twitter, LinkedIn, etc.