Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv Computation and Language
6d ago

Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

arXiv:2609.37837v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly el...

By Haotian Deng, Wenbin Xing, Gang Xu, Tao He, Jinkai Zheng, Chun Li, Zheng Zhu, Ming Li
arXiv Computer Vision
6d ago

HelixWorld: A Real-time Interactive Audio-Visual World Model

arXiv:2609.38123v1 Announce Type: new Abstract: World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models...

By Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang, Haoyuan Huang, Kam Man Wu, Pengjun Fang, Hongyu Liu, Chenyang Qi, Lin Wang, Ruibin Yuan, Weijia Chen, Fangneng Zhan, Qifeng Chen, Wei Xue, Yike Guo
arXiv AI
6d ago

PILLAR: Private Inverted-Index Lexical Lookup for Augmented Retrieval

arXiv:2609.36326v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system bas...

By Truong Son Nguyen (Arizona State University), Daniel Blackley (George Mason University), Ni Trieu (Arizona State University), Evgenios M. Kornaropoulos (George Mason University)