arXiv:2609.36392v1 Announce Type: new
Abstract: In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis...
By Yuyan Chen
arXiv:2609.37588v1 Announce Type: new
Abstract: Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the...
By T. Duy Nguyen-Hien, Yee Whye Teh, Wee Sun Lee, Tan Zhi-Xuan
arXiv:2609.35860v1 Announce Type: cross
Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors a...
By Pranav Darshan, Pranav A, Sravan Karthick T, Minal Moharir, Ivan P. Yamshchikov
arXiv:2605.01418v2 Announce Type: replace
Abstract: Time-series data are inherently multiscale, spanning diverse temporal granularities from coarse trends to fine-scale dynamics. However, existing ti...
By Seokhyun Lee, Jaeho Kim, Changjun Oh, Mihaela van der Schaar, Changhee Lee
arXiv:2603.04410v3 Announce Type: replace-cross
Abstract: While different stakeholders are trying to leverage Arabic Language Models (ALMs), safety alignment in ALMs remains largely underexplored, hi...
By Omar Abdelnasser, Fatemah Alharbi, Khaled Khasawneh, Ihsen Alouani, Mohammed E. Fouda
arXiv:2609.36302v1 Announce Type: new
Abstract: While foundation models have revolutionized natural language processing and computer vision by leveraging universal vocabularies, Graph Machine Learnin...
By Ben Finkelshtein, Andr\'{e} Linhares, Petar Veli\v{c}kovi\'{c}, Bryan Perozzi, Mikhail Galkin
arXiv:2609.36724v1 Announce Type: new
Abstract: Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gr...
By Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong
arXiv:2609.37604v1 Announce Type: new
Abstract: Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges sp...
By Yuxiang Yao, Zijun Zhao
arXiv:2609.35881v1 Announce Type: cross
Abstract: Can a genome model retain a detectable record of the synthetic sequences used to train it? We study watermark inheritance through distillation with G...
By Guang Yang, Fengchen Liu
arXiv:2609.35932v1 Announce Type: cross
Abstract: Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged templat...
By Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
arXiv:2609.36798v1 Announce Type: cross
Abstract: Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training para...
By Yueran Ma, Ronghao Lin
arXiv:2609.37882v1 Announce Type: cross
Abstract: Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other Afri...
By Bhanu Prakash Vangala, Sowmya Guda, Navya Vangala
The paper investigates how the choice of graph tokenization affects transformer expressivity. It analyzes three tokenization families—spectral, random‑walk, and adjacency—showing that each induces different depth requirements and that some tokenizations are inherently lossy or ill‑conditioned for certain tasks. The authors prove lower bounds and impossibility results for converting between tokenizations and validate these findings with experiments on synthetic and real‑world data.
By Maya Bechler-Speicher, Gilad Yehudai, Gil Harari, Clayton Sanford, Amir Globerson, Joan Bruna
arXiv:2603.28773v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) frequently generate confident yet factually incorrect content when used for language generation (a phenomenon of...
By Dobrik Georgiev, Kheeran K. Naidu, Alberto Cattaneo, Federico Monti, Carlo Luschi, Daniel Justus
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --...
The article titled "Language Models for Text Classification: From Bag-of-Words to Jev" offers a visual guide that explores various neural network architectures—including RNNs, CNNs, and Transformers—alongside calibration techniques. It includes hands‑on experiments that compare the accuracy and efficiency of these models for text classification tasks.
By Sebastian Raschka, PhD
Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a...
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover de...
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation qual...
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradien...