arXiv AI By Obed Junias, Maria Leonor Pacheco

From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

Read the original on arXiv AI →

arXiv:2608. 12836v1 Announce Type: cross Abstract: Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 14

PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains

arXiv:2508. 21787v3 Announce Type: replace-cross Abstract: Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate solutions and selecting the one with the highest reward.

By Joshua Ong Jun Leang, Zheng Zhao, Aryo Pradipta Gema, Sohee Yang, Wai-Chung Kwan, Xuanli He, Wenda Li, Pasquale Minervini, Eleonora Giunchiglia, Shay B. Cohen