Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,725 stories · RSS feed

arXiv Computation and Language
5d ago

Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification

arXiv:2609.38812v1 Announce Type: new Abstract: Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments....

By Yingfeng Luo, Shaowei Wei, Daixin Wang, Dingyang Lin, Kaiyan Chang, Weiqiao Shan, Tong Zheng, Zhiqiang Zhang, Jingbo Zhu, Tong Xiao
arXiv Computer Vision
5d ago

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

arXiv:2609.34330v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...

By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang