Skip to content

Mixture of experts

Peak · emerged January 2017 · high July 2026 · +7% on the month before · 16 mentions on Trending in 14 days

Models split into many expert blocks with only a few active per token, so they grow in size without growing in cost. Named by Shazeer et al., Google Brain: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Jan 23, 2017). Also called MoE, Sparse models, Sparse mixture of experts.

Index 0 to 100, monthly
038762017201820192020202120222023202420252026Origin

Signals

SignalReads as
Hacker News stories, last 12 months93 stories, +182% on the 12 months beforeAdoption
arXiv papers, last 12 months1,835 papers, +93% on the 12 months beforeAdoption
Wikipedia views, August 20268,562, -21% on a year beforeDecline
Product releases naming it13 since Sep 1, 2025Adoption
GitHub repositories2,012Adoption

Companies

Releases

13 releases since Sep 1, 2025
Name
Transformers Release 5.17.0Hugging Face · Sep 10, 2026 2 weeks ago
Announcing Cohere's North Small TranslateCohere · Sep 9, 2026 2 weeks ago
Zhipu AI GLM 5.3 now available as a Databricks-hosted modelDatabricks · Aug 29, 2026 3 weeks ago
Hy4 Preview now available on AI GatewayVercel · Aug 28, 2026 4 weeks ago
Workers AI - Z.ai GLM-5.3 Flash now available on Workers AICloudflare · Aug 26, 2026 4 weeks ago
Ling 3.0 Flash is now available on AI GatewayVercel · Jul 23, 2026 2 months ago
Laguna S 2.1 is now available on AI GatewayVercel · Jul 21, 2026 2 months ago
Thinking Machine Labs Inkling now available as a Databricks-hosted model in Public PreviewDatabricks · Jul 15, 2026 2 months ago

Models

Papers

Name
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionDeepSeek · Sep 17, 2026 8 days ago
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training StabilityAlibaba (Qwen) · Aug 31, 2026 3 weeks ago

GitHub projects

2,012 repositories
NameStars
deepseek-ai/DeepSeek-VL2DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding · Dec 13, 2024 1 year ago5,375
deepseek-ai/DeepSeek-V2DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model · Apr 22, 2024 2 years ago5,041
PKU-YuanGroup/MoE-LLaVA【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models · Dec 14, 2023 2 years ago2,324
deepseek-ai/DeepSeek-MoEDeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models · Jan 2, 2024 2 years ago1,976
XueFuzhao/OpenMoEA family of open-sourced Mixture-of-Experts (MoE) Large Language Models · Aug 8, 2023 3 years ago1,699
XueFuzhao/awesome-mixture-of-expertsA collection of AWESOME things about mixture-of-experts · Mar 30, 2022 4 years ago1,287
MovementStage
Open-weight modelsModels whose trained weights anyone can download, run and fine-tune: Llama, Qwen, DeepSeek, Mistral and more.Rising
Inference chipsAccelerators built to serve models cheaply and fast rather than to train them: TPUs, LPUs, wafer-scale and custom silicon.Peak
Reasoning modelsModels trained with reinforcement learning to think step by step, spending more compute at answer time.Plateau

History

DateWhat changed
Sep 24, 2026Tracking started: emerged January 2017, peak, 3 other names recorded

Sources: Curve: Hacker News story titles, Wikipedia pageviews and arXiv papers by month, refreshed Sep 24, 2026. Related entities from the fru.dev sites' public APIs, matched by name. How stages work.

Trending terms by email

Monday mornings: the week's top trending terms in data, tech and AI.

Double opt-in. Unsubscribe any time.