arXiv AI By Kang Chen, Sihan Zhao, Yixin Cao, Yugang Jiang

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

Read the original on arXiv AI →

arXiv:2608. 17638v1 Announce Type: new Abstract: What a reasoning model writes is only a partial record of the process that produces it.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.