arXiv Machine Learning By Tailia Malloy, Prateek Kumar Rajput, Serge Lionel Nikiema, Cleotilde Gonzalez, Tegawend\'e F. Bissyand\'e

Metacognitive Reasoning in Energy Based Models using Instance Based Learning Theory

Read the original on arXiv Machine Learning →

The paper introduces MERITED, a framework that combines Instance-Based Learning Theory (IBLT) with Energy Based Models (EBMs) to enable metacognitive reasoning about computational effort. It allows an EBM to dynamically allocate resources based on uncertainty estimates, addressing limitations of large language models that cannot predict uncertainty before responding. The authors present a 191‑million‑parameter reasoning EBM and demonstrate how MERITED uses an IBL model for efficient, uncertainty‑driven compute allocation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 10

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents

arXiv:2510. 22052v2 Announce Type: replace Abstract: The field of artificial intelligence (AI) has taken a tight hold on broad aspects of society, industry, business, and governance in ways that dictate the prosperity and might of the world's economies.

By Abhijit Chatterjee, Niraj K. Jha, Jonathan D. Cohen, Thomas L. Griffiths, Hongjing Lu, Diana Marculescu, Ashiqur Rasul, Wenrui Xu, Keshab K. Parhi
arXiv AI
Sep 18

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

When2Think introduces a post‑training framework that dynamically allocates reasoning depth in Large Reasoning Models based on instance difficulty. The method uses Instance‑level Difficulty‑Aware Control (IDAC) to shape rewards with pre‑computed accuracy and token usage statistics, enabling stable, critic‑free optimization without learned reward models. Experiments on mathematical benchmarks show that When2Think improves accuracy‑efficiency trade‑offs, achieving higher Pass@3 scores while reducing token usage compared to baseline models.

By Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak
arXiv AI
Jun 4

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

arXiv:2606. 04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset.

By Guangcheng Zhu, Shenzhi Yang, Haobo Wang, Xing Zheng, Yingfan MA, Xuening Feng, Zhongqi Chen, Bowen Song, Weiqiang Wang, Gang Chen