A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence c...
arXiv:2511. 04919v3 Announce Type: replace Abstract: Processing long documents with large language models (LLMs) is expensive: a single query over a 100K-token document can cost from tens of cents to over a dollar in API fees, depending on the model, and memory grows linearly with context length.
By Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi, Sheikh Nazib Ahmed
A$^2$Safe is a framework for safe and effective Visual Question Answering that aligns counterfactual evidence with adaptive agent collaboration. It uses a Grounded Safety Evidence Board to make safety decisions explicit, enforcing invariance to safety‑irrelevant changes while allowing appropriate transitions when risk‑critical evidence changes. The system achieves a 95.72 SIUO safety score, reduces benign refusals on MOSSBench to 14.67%, and maintains a 78.34 average VQA score with 27.8% token overhead.
By Quanxing Xu, Ling Zhou, Xian Zhong, Jinyu Tian, Xiaohua Huang, Rubing Huang, Chia-Wen Lin
The paper introduces the ACE framework, which uses spatially explicit inspection to guide embodied exploration. By combining evidence‑grounded perception with exposure‑informed movement, ACE provides a spatially resolved decision paradigm that improves cue assessment and movement direction. Experiments show ACE boosts navigation task success by 18.0% and exploration efficiency by 10.3% over previous baselines.
By Wenbin Wang, Xiang Bai, Yizhao Wang, Hang Sun, Dong Ren, Jie Qin, Qingquan Li, Bing Wang
arXiv:2609.22097v1 Announce Type: cross
Abstract: The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advan...
By Junpeng Wang, Yuzhong Chen, Menghai Pan, Uday Singh Saini, Yiwei Cai
arXiv:2609.23703v1 Announce Type: cross
Abstract: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrat...
By Kemal Kirtac
arXiv:2609.24569v1 Announce Type: cross
Abstract: Over the past decade, a growing body of research has shown that $\gamma$-weak submodularity broadly arises in numerous subset selection tasks, includ...
By Shi Fu, Youming Qiao, Dacheng Tao, Zongqi Wan, Qixin Zhang
arXiv:2502.01562v3 Announce Type: replace
Abstract: As the general capabilities of artificial intelligence (AI) agents continue to evolve, their ability to learn to master multiple complex tasks thro...
By Minttu Alakuijala, Ya Gao, Georgy Ananov, Samuel Kaski, Pekka Marttinen, Alexander Ilin, Harri Valpola
arXiv:2511.04473v3 Announce Type: replace
Abstract: Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs. While various...
By Alberto Cattaneo, Carlo Luschi, Daniel Justus
arXiv:2605.08992v2 Announce Type: replace
Abstract: Federated learning (FL) is increasingly used to fine-tune foundation models (FMs) on distributed private data. The community largely assumes that l...
By Kiran Naseer, Umar Shoaib
arXiv:2609.22114v1 Announce Type: new
Abstract: Context compression is widely proposed as a way to cut the token bill of LLM coding agents, and public benchmarks report that aggressive compression pr...
By Luzhuo Chen, Jiayu Shi
arXiv:2609.22125v1 Announce Type: new
Abstract: Standard tokenizers used in large language models produce malformed text when applied to Brahmic scripts. They are a family of abugidas, writing system...
By Sai Hemanth Kapila, Rakshika Bagavathy
arXiv:2609.22210v1 Announce Type: new
Abstract: SALSA (Semi-Autonomous Literature Summarization Assistant) is an open- source, human-in-the-loop platform for extracting structured scientific datasets...
By William Schertzer, Sonakshi Gupta, Rampi Ramprasad
arXiv:2609.22213v1 Announce Type: new
Abstract: Temporal Knowledge Graph Question Answering (TKGQA) requires answer inference from evidence that is both structurally valid and temporally admissible....
By Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Dongjin Yu, Yu Wang
arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv:2609.22463v1 Announce Type: new
Abstract: Sleep monitoring using wearable data has shown promise for personal health, yet large language model (LLM)-based summarization and question answering r...
By Yusheng Tan, Running Zhao, Sofia Angel, Ninghui Hao, Ash Arian, Nikita N. Dulin, Jay Lin, Ou Zhu, Faiza Shaik, Xinxing Yang, Bonnie W. Leung, Katie Roster, Arlene Ruiz de Luzuriaga, Kenneth Lee, Alejandra Lastra, Habibul Ahsan, Guihong Wan
arXiv:2609.22796v1 Announce Type: new
Abstract: Dialectal Arabic machine translation (MT) remains challenging despite recent progress in Arabic language technologies, particularly because effective t...
By Abdellah El Mekki, AbdelRahim A. Elmadany, Samar M. Magdy, Saad Ezzini, Mo El-Haj, Mustafa Jarrar, Zaid Alyafeai, Bernard Ghanem, Muhammad Abdul-Mageed
arXiv:2609.23056v1 Announce Type: new
Abstract: Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear...
By Kai-Hsin Chen, Wei-Yu Chen, Xuanjun Chen, Jyh-Shing Roger Jang
arXiv:2609.23726v1 Announce Type: new
Abstract: Large language models have shown strong performance across a range of legal tasks, but existing benchmarks rarely evaluate the ability to take and defe...
By Jiakang Xu, Wantong Huo, Udom Silparcha, Jonathan H. Chan
arXiv:2609.23955v1 Announce Type: new
Abstract: Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representa...
By Mai Mohamed Eida, Ryan Dolan, Paul de Nijs, Jonathan Dunn