Writing
September 21, 2026 · 6 min read

IntBMoE: what if a mixture-of-experts model didn't have to pick two out of three

IntBMoE decouples participation, execution, and materialization in Mixture-of-Experts by composing experts with a hypernetwork before routing sparsely to a fixed codebook of blocks — and it's already running in AMap's recommender at scale.

mixture-of-expertsmodel-architecturerecommendation-systemsproduction-mlefficient-inference

Decoupling how many experts contribute from how many are computed in a mixture-of-experts network

Every Mixture-of-Experts design I've worked with or read about eventually runs into the same wall: you get to pick two of three good properties, not all three. A new paper, IntBMoE (Cheng, Xu, Liu, Liu, and Chu), names the three properties precisely enough that the tradeoff becomes obvious in hindsight, and then proposes an architecture that avoids picking at all. What makes it worth a post rather than a shrug is the last line of the abstract: it's already deployed in AMap's generative recommendation system, serving hundreds of millions of users inside a 60ms latency budget, with a measured 2.4% relative lift in UVCTR from an online A/B test.

Three numbers, not one

The paper's contribution starts with vocabulary. For a single token passing through an MoE layer, there are three separate quantities that most of us collapse into "how many experts are active":

  • Participation — how many experts actually contribute knowledge to the token's output.
  • Execution — how many experts are actually computed (your compute cost).
  • Materialization — how many expert-sized parameter sets have to exist in memory at all (your memory cost).

Standard sparse routing (top-k MoE) keeps execution and materialization low, which is exactly why it's the default at scale: you route each token to a handful of experts and only pay for those. But participation drops with it — most of the model's knowledge never touches that token. Dense output-mixing (compute every expert, blend the outputs) restores full participation, but execution now scales with the total number of experts, which is the thing sparse routing was invented to avoid. Parameter-merging schemes (blend expert weights into one expert before running it) hold execution at a single expert, but the number of distinct merged parameter sets you need to store grows with the number of routing decisions you make — materialization balloons instead.

Three axes, three failure modes, and no existing design sets all three independently. That's the setup.

Composing before routing, not instead of routing

IntBMoE's move is to stop tying "which experts contribute" and "which experts get computed" to the same decision. It introduces a small, learned codebook of blocks — a fixed, bounded set of composed-expert slots, independent of how many tokens or routing decisions occur downstream. At each internal layer, a lightweight hypernetwork takes every expert basis in that layer's pool and merges them into one composed expert per codebook entry, ahead of the router seeing any token.

That single design choice is what decouples the three quantities:

  • Participation is full, because every composed block already draws on the entire expert pool — the hypernetwork did the mixing before any token arrived, so there's no per-token "only a few experts saw this" gap.
  • Execution stays sparse, because the router still only sends each token to a few blocks, not the whole codebook.
  • Materialization is bounded, because the number of composed experts that ever need to exist is fixed by the codebook size, not by the number of tokens or routing decisions made at inference time.

The image below is the shape of that argument: sparse routing, dense mixing, and parameter merging each trade away one of the three properties, and IntBMoE's codebook-plus-hypernetwork pipeline is what lets it hold all three.

Comparison of sparse routing, dense mixing, parameter merging, and IntBMoE across participation, execution, and materialization, plus the codebook-hypernetwork-router pipeline

On top of that, the paper adds Dual-Path Residual Gating (DPRG): two independently composed paths through the block codebook, coupled by multiplicative gating rather than a simple sum. Multiplicative gating lets one path modulate the other rather than just averaging with it, which is a cheap way to add expressivity without adding a third path or more execution cost.

Where it was tested, and where it runs

The experimental section covers image classification (their primary benchmark, showing consistent gains over both sparse and dense MoE baselines), then extends to language modeling and sequential recommendation to check that the idea isn't vision-specific. That's a reasonable generalization test, and the fact that a block-conditioning idea developed for vision transfers to sequence models and recommendation towers is a decent signal that the underlying mechanism — decoupling composition from routing — isn't tied to one architecture family.

The part I'd weight most heavily, though, is the production claim: IntBMoE is fully deployed in AMap's generative recommendation system, serving hundreds of millions of users, inside a 60ms latency budget, and it produced a 2.4% relative UVCTR gain in online A/B testing. Academic MoE papers routinely stop at offline benchmark tables. A recsys team putting this into a live 60ms serving path and reporting a real A/B result is a different order of evidence — it means the sparse-execution promise held up against an actual latency SLO, not just a FLOPs count on paper.

Why I'd pay attention to this

I lead production AI systems where inference cost is a line item, not an abstraction, and the participation/execution/materialization framing is genuinely useful even outside this specific architecture. It gives you a vocabulary to diagnose why a given MoE variant is expensive: is it expensive because you're computing too much (execution), storing too much (materialization), or because it's cheap but shallow (participation)? Most MoE postmortems I've seen conflate these, which makes it harder to reason about which lever to pull.

Whether IntBMoE's specific hypernetwork-composition trick becomes standard practice is a separate question from whether the framing is right, and I think the framing is right regardless. The AMap deployment matters because it's the difference between "this decouples the three quantities on paper" and "this decouples the three quantities under a 60ms SLO with real user traffic." That's the bar production ML work has to clear, and it's rare to see a paper clear it in the same document that introduces the idea.

The authors have released code, linked from the arXiv abstract page, for anyone who wants to check whether the hypernetwork composition step adds meaningful overhead in their own stack before the router even runs.

References
  1. 01IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts (arXiv:2609.21346)