Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv AI
6d ago

KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora

arXiv:2609.37673v1 Announce Type: new Abstract: Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to tak...

By Changmian Wang, Yuchao Ma, Xuchao Lu, Chen Zhang, Ping Sun, Jiazheng Wang, Shan Wang, Xuanwen Chen, Yihe Sun, Ziyu Lu, Jianqiang Huang, Hongzhi Li, Ziqing Xia, Kaihua Tang, Xian-Sheng Hua, Qinghua Zheng
arXiv AI
6d ago

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

arXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...

By Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin