← Back to all news
arXiv Computer Vision September 22, 2026 By Weihan Cai, Hao Tan, Xinping Gao, Shibiao Xu, Jun Wan

PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • fine-tuning
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 9

An Effective Router for Vision-Language Model Selection

arXiv:2606. 08970v1 Announce Type: new Abstract: Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making it difficult for users to select the most appropriate one among numerous VLM candidates.

By Can Wang, Shengwei Wang, Bolin Zhang, Zhiying Tu, Dianhui Chu
llmsmultimodal
More like this →
arXiv AI
4d ago

Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

arXiv:2609.37362v1 Announce Type: new Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--eff...

By Guannan Lai, Han-Jia Ye
llms
More like this →
arXiv AI
Jul 15

ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing

arXiv:2603. 24787v2 Announce Type: replace Abstract: Routing has emerged as a promising strategy for balancing performance and cost in large language model (LLM) systems that combine lightweight models with powerful but expensive large models.

By Yaopei Zeng, Congchao Wang, Blake JianHang Chen, Lu Lin
llmsfine-tuningmultimodal
More like this →
arXiv AI
Sep 15

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

arXiv:2609.15131v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most method...

By Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang
llmsreinforcement-learningmultimodal
More like this →
arXiv AI
Jul 8

Self-Routing: Parameter-Free Expert Routing from Hidden States

arXiv:2604. 00421v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to expert assignments.

By Jama Hussein Mohamud, Drew Wagner, Mirco Ravanelli
More like this →
arXiv Computer Vision
Sep 10

Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs

arXiv:2609.10346v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existin...

By Haiji Liang, Pengfei Zhou, Zhenglin Wan, Wei Wang, Yang You, Wangbo Zhao
llmsefficiencymultimodalbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea