Hugging Face Trending Papers

Cascaded Batch Prompting

Cascaded Batch Prompting introduces a two‑stage method to improve large language model inference by separating complex reasoning from symbol grounding, addressing the unpredictability of traditional batch prompting. Experiments on multiple‑choice question answering and natural language inference show that this approach outperforms single prompting while scaling speed with batch size, achieving a new state‑of‑the‑art balance between accuracy and efficiency.

arXiv Computation and Language
Aug 28

Cascaded Batch Prompting

Cascaded Batch Prompting introduces a two‑stage method that separates complex reasoning from symbol grounding to address the unpredictability of conventional batch prompting. Experiments on multiple‑choice question answering and natural language inference show that this approach outperforms standard single prompting while maintaining a speedup proportional to batch size. The technique establishes a new state‑of‑the‑art position on the Pareto frontier for efficiency and performance.

By Sho Hoshino, Peinan Zhang
arXiv Computation and Language
Sep 22

SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine

arXiv:2410.17021v2 Announce Type: replace Abstract: Large Language Models with chain-of-thought prompting, such as OpenAI-o1, have shown impressive capabilities in natural language inference tasks. H...

By Xiaochen Wang, Liang Chen, Reza Haf Zhe Yang, Yiru Wang, Xiangdi Meng, Kunhao Pan, Zhifang Sui, Junqing He
arXiv Computation and Language
Sep 25

Likelihood Ranking doesn't Scale Like Prompting in LLMs

The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.

By Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci
arXiv AI
Sep 7

Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

The paper surveys efficient reasoning in large language models, contrasting fast intuitive (System 1) and slow deep (System 2) reasoning. It analyzes why System 2 is computationally costly yet more accurate, and why System 1 is efficient but less effective. The survey covers causes of inefficiency, patterns of reasoning behavior, and potential solutions to balance performance and computational budgets, offering actionable insights and an open‑source repository for ongoing research.

By Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, Kam-Fai Wong
arXiv Computation and Language
Sep 1

HiVe: Beyond Static Prompts for Multitask Learning via Hierarchy-based Vertical Mixture-of-Experts

HiVe is a prompt‑tuning framework that builds a hierarchy of prompts by exploiting inter‑task relationships during training. It uses a vertical mixture‑of‑experts (V‑MoE) at inference to compose prompts at the level of specialization needed for each input, allowing input‑dependent prompt adaptation. Experiments demonstrate that HiVe consistently outperforms strong prompt‑tuning baselines across diverse tasks.

By HyeonJik Bae, Minyeol Kim, Susik Yoon