Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
Researchers have explored whether large neural networks can be designed to cache frequently used parts of their architecture, reducing memory bandwidth bottlenecks. They trained a mixture-of-experts model to prioritize locality and domain router losses, but found that this approach fails to meet the desired perplexity threshold while achieving significant cache miss reduction. The study also evaluated training-free cache-aware rerouting stacks and domain-primed prefetching me
Researchers have explored whether large neural networks can be designed to cache frequently used parts of their architecture, reducing memory bandwidth bottlenecks. They trained a mixture-of-experts model to prioritize locality and domain router losses, but found that this approach fails to meet the desired perplexity threshold while achieving significant cache miss reduction. The study also evaluated training-free cache-aware rerouting stacks and domain-primed prefetching methods, which showed promise in reducing misses at lower perplexity costs.
---
Why it matters: This research matters because it sheds light on the limitations of current large-scale neural network architectures and their memory bandwidth requirements. Understanding how to optimize these models for edge devices is crucial for widespread adoption in applications such as voice assistants and translation services.
Source: https://arxiv.org/abs/2608.18261
This article was originally published at: https://arxiv.org/abs/2608.18261