arXiv:2608. 06377v1 Announce Type: cross Abstract: Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong.
By Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
arXiv:2608. 06283v1 Announce Type: new Abstract: We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex.
By Iosif Lytras, Nikolaos Makras, Sotirios Sabanis
arXiv:2608. 05611v1 Announce Type: cross Abstract: Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy.
By Guanyu Wang, Zidi Zhang, Xu Chu
arXiv:2608. 05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings.
By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
arXiv:2608. 05802v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored.
By Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
arXiv:2608. 06020v1 Announce Type: new Abstract: Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes.
By Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong
arXiv:2601. 21124v2 Announce Type: replace-cross Abstract: Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI.
By Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
arXiv:2604. 09297v3 Announce Type: replace-cross Abstract: Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as static, hand-crafted assets, or evolved on pass rate alone.
By Jingzhi Gong, Ruizhen Gu, Zhiwei Fei, Yazhuo Cao, Lukas Twist, Alina Geiger, Shuo Han, Dominik Sobania, Federica Sarro, Jie M. Zhang
arXiv:2608. 06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations.
By Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li
arXiv:2608. 05232v1 Announce Type: cross Abstract: The work of Tang et.
By Patrizia Kaye
arXiv:2608. 05161v1 Announce Type: cross Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining remains an unsolved practical challenge.
By Josh McGiff, Salma Mekaoui, Robert Shanahan, Nikola S. Nikolov
arXiv:2608. 06310v1 Announce Type: new Abstract: Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models.
By Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang, Tongran Liu, Jingbo Zhu
arXiv:2608. 05367v1 Announce Type: new Abstract: Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making.
By Zonghao Yang
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa
arXiv:2608. 05371v1 Announce Type: new Abstract: World models learn latent states that summarize interaction histories, evolve over time, and support prediction, simulation, or planning.
By Hailong Jiang, Emran Hossain, Feng Yu, Jianfeng Zhu, Guilin Zhang, Wulan Guo
arXiv:2608. 06004v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) are currently the best approach to tabular prediction problems.
By Christian Kl\"otergens, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme, Tom Hanika
arXiv:2608. 05790v1 Announce Type: new Abstract: General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments.
By Jiacheng Wei, Zhaoxin Fan, Xin Wen, Yuqin Lan, Dongrun Li, Wenjun Wu, Faguo Wu, Xiao Zhang
arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.
By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
arXiv:2608. 05872v1 Announce Type: cross Abstract: Standard Large Language Models (LLMs) execute layers sequentially.
By Pawe{\l} Batorski, Abtin Pourhadi, Akylgali Aitaza, Przemys{\l}aw Spurek, Paul Swoboda
arXiv:2608. 06197v1 Announce Type: new Abstract: Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground.
By Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu