arXiv AI

CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions

arXiv:2607. 08774v1 Announce Type: new Abstract: Reliability in large language model (LLM) systems is typically framed as a function of model capability.

arXiv Computation and Language
Aug 28

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

CritICL is an inference-time framework that enhances reasoning in large language models by using failure patterns from weaker models as critique-based in-context examples. It offers two variants: CritICL-dynamic, which predicts input-specific failure modes, and CritICL-static, which applies a global failure mode profile. Experiments show that CritICL outperforms standard in-context learning and rivals test-time scaling methods while using fewer generations and lower token costs.

By Yufan Wu, Yinghui He, Zhengyi Hu, Lang Wei, Ruichen Li, Qifan Yang, Ting Zhu
arXiv Machine Learning
Jun 2

Are Large Reasoning Models Interruptible?

arXiv:2510. 11713v4 Announce Type: replace-cross Abstract: Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments.

By Tsung-Han Wu, Mihran Miroyan, David M. Chan, Trevor Darrell, Narges Norouzi, Joseph E. Gonzalez
arXiv AI
Sep 4

EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering

EasySteer is a unified framework for high‑performance, extensible large‑language‑model steering built on vLLM. It offers a modular architecture with pluggable interfaces for analysis‑based and learning‑based methods, fine‑grained parameter control, pre‑computed steering vectors for eight application domains, and an interactive demo system. Integrated with vLLM’s optimized inference engine, EasySteer delivers a 10.8–22.3× speedup over existing frameworks and demonstrates effectiveness in overthinking mitigation, hallucination reduction, and other key applications.

By Haolei Xu, Xinyu Mei, Yuchen Yan, Rui Zhou, Wenqi Zhang, Weiming Lu, Yueting Zhuang, Yongliang Shen
arXiv AI
Sep 18

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

PetriBench is a compact, fully self‑contained, and scalable benchmark that evaluates large language model (LLM) reasoning over dynamic state spaces using Petri nets. It organizes reasoning into four task families with Easy, Medium, and Hard levels, each generated by increasing structural complexity and evaluated against exact ground truth. Experiments across proprietary and open‑weight models show that accuracy consistently drops with difficulty, revealing distinct task‑specific capability profiles, while test‑time compute and procedural generation affect performance differently across tasks.

By Pyrros Koussios, Benjamin J\"ager, John Hua Yao, Ajay Sridhar, Violet Xiang, Chenhao Li