Mixture of Experts (MoE)
An architecture that routes each token to a few specialized sub-networks, boosting capacity without proportional cost.
An architecture that routes each token to a few specialized sub-networks, boosting capacity without proportional cost.