Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Abstract
SharpMoE addresses routing inefficiencies in diffusion models by using clean latent features to guide salient token identification and employs trajectory routing loss for precise compute allocation during multi-step denoising.
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocating computational resources across diverse tokens to improve efficiency and performance. However, we identify a routing assignment problem in existing diffusion MoE frameworks: the router fails to accurately allocate more computational resources to salient tokens. Our analysis attributes this failure to the router's reliance on noise-corrupted latent features throughout the denoising process. Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. By bypassing the noise-distorted inputs, SharpMoE provides the router with clear saliency guidance, enabling the identification of salient tokens even in high-noise stages. Furthermore, we introduce a trajectory routing loss to constrain the compute allocation throughout the multi-step denoising trajectory, ensuring precise resource allocation along the generation rollout. Extensive experiments demonstrate that SharpMoE serves as a versatile, plug-and-play solution that further enhances the pretrained, converged MoE models, achieving state-of-the-art performance in visual generation.
Community
A Diffusion MoE Framework with Saliency-Harnessing Accurate Routing
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion (2026)
- FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion (2026)
- AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers (2026)
- Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers (2026)
- EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation (2026)
- Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts (2026)
- Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2606.26938 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper