arXiv Computation and Language By Xiaotian Zhang (Trooly.AI)

Beyond Depth and Width: The Information-Slack Dilemma in Streaming Test-Time Compute

Read the original on arXiv Computation and Language →

The paper discusses how the same computational task can require different reasoning strategies depending on the order in which evidence arrives, introducing the concept of an "information‑slack dilemma." It argues that early computation may be useful only if its benefits outweigh the costs of later verification, invalidation, and recovery, and proposes a research agenda focused on selective recovery and predictive policies. The authors emphasize evaluating these approaches by separating early‑execution effects, deployment value versus full‑input alternatives, and the added value of predictive policies while considering shared‑resource costs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
2d ago

Intelligence Under Time Constraints: Rethinking Test-Time Compute

The paper "Intelligence Under Time Constraints: Rethinking Test-Time Compute" examines how to decide not only how much computation to perform but also when to start it in streaming interactions where evidence arrives incrementally and may change. It introduces the information‑slack dilemma, describing the trade‑off between early computation that has more time to finish but relies on incomplete evidence, and waiting for more information that reduces computational slack. The authors propose a research agenda focused on selective recovery under controlled evidence revisions, and suggest evaluation metrics that separate early‑execution effects, deployment value versus full‑input alternatives, and the added value of predictive policies while accounting for shared‑resource costs, aiming for trustworthy, on‑time responses within a declared resource envelope.

By Xiaotian Zhang (Trooly.AI)
arXiv AI
Aug 5

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

arXiv:2608. 04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute.

By Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary