arXiv:2608. 07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it shouldn't.
By Aditya Katkar, Om Karkele, Kartik Mandhane, Manisha More, Yash Kashid
arXiv:2608. 07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens.
By MY Pitsane, Hope Mogale
arXiv:2608. 07363v1 Announce Type: new Abstract: Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility bursts, structural shifts, and nonlinear oscillatory behaviors.
By Junkai Lin, Siqi Hou, Raymond Lee
arXiv:2608. 07364v1 Announce Type: new Abstract: Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the Curriculum as Code paradigm, integrating Generative AI with LaTeX and Python to automate the creation of reproducible, visually consistent, and technically precise materials for STEM education.
By Henrique Mohallem Paiva
arXiv:2608. 07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.
By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
arXiv:2608. 07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated.
By Ruochen Jin, Zhanliang Wang, Zongyu Dai, Jiancong Xiao, Bojian Hou
arXiv:2608. 07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited.
By Vasanth Iyer
arXiv:2608. 07367v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern.
By Maria-Louisa Wightman, Guillaume Bied, Tijl De Bie
arXiv:2608. 07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities.
By Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine
arXiv:2608. 06471v1 Announce Type: cross Abstract: Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software.
By Amine Lbath, Manan Suri, Aurelien Delaitre, Vadim Okun, Massih-Reza Amini, Ram D. Sriram, Dinesh Manocha
arXiv:2602. 03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget.
By Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
arXiv:2602. 20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output.
By Peter Hase, Christopher Potts
arXiv:2603. 21607v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so robust uncertainty quantification (UQ) remains essential.
By Alexandra Bazarova, Andrei Volodichev, Daria Kotova, Alexey Zaytsev
arXiv:2604. 25220v2 Announce Type: replace Abstract: Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in journalism, education, and public communication.
By Ridwan Mahbub, Syem Aziz, Mizanur Rahman, Mahir Ahmed, Shadikur Rahman, Shafiq Joty, Enamul Hoque
arXiv:2604. 16009v2 Announce Type: replace Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or conflicting evidence.
By Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane
arXiv:2502. 20295v3 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task.
By Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin, Xiaowen Dong
arXiv:2507. 08390v5 Announce Type: replace Abstract: Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-scale training.
By Meihua Dang, Jiaqi Han, Minkai Xu, Kai Xu, Akash Srivastava, Stefano Ermon
arXiv:2608. 07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims.
By Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou
arXiv:2601. 19897v2 Announce Type: replace Abstract: Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models.
By Idan Shenfeld, Mehul Damani, Jonas H\"ubotter, Pulkit Agrawal
arXiv:2608. 06417v1 Announce Type: new Abstract: The proliferation of misinformation online has driven demand for scalable detection systems.
By Pedro Barcelos, Ot\'avio Parraga, Marcelo M. Mussi, Lucas M. Fraga, Lucas S. Kupssinsk\"u, Rodrigo C. Barros