arXiv:2607.08152v2 Announce Type: replace-cross
Abstract: Predicting comprehension from eye movements could support adaptive reading interfaces. We present LEXIC, a compact recurrent model that predi...
By Sumin Lee, Kyeonghun Kim, Subeen Lee, Jiwon Yang, Hyunsu Go, Eunseob Choi, Ken Ying-Kai Liao, Nam-Joon Kim
The study examines how reader proficiency influences the relationship between layer-wise surprisal from large language models (LLMs) and eye-tracking gaze measures. Using the MECO L2 corpus, researchers compared high- and low-proficiency readers on first-pass gaze duration (FPGD) and total gaze duration (TGD), finding that lower-proficiency readers exhibit deeper Predictive Depth for FPGD, while TGD shows deeper Predictive Depth across both groups. The results suggest that the distribution of predictive power across LLM layers relates to the timing and breadth of reading processes and varies with reader proficiency.
By Akio Hayakawa, Horacio Saggion
arXiv:2608.30583v1 Announce Type: new
Abstract: Standard language proficiency tests rely on linguistic tasks such as vocabulary, grammar and reading comprehension quizzes. An alternative, cognitively...
By Shachar Frenkel, Ido Falah, Omer Shubi, Yevgeni Berzak
arXiv:2606. 28876v2 Announce Type: replace-cross Abstract: We study memory-managed long-context attention: explicit bounded memory with a learned query-independent writer, lifecycle control, query-aware reading, calibrated sparse fallback, and frozen-LLM generation from raw evidence.
By Junyi Zou, Avrova Donz
The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.
By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
The report introduces A.X K2, a 688‑parameter Mixture‑of‑Experts language model designed for agentic applications. Trained on 8.5 trillion tokens, it surpasses its predecessor A.X K1 by over 30 percentage points on several benchmarks, thanks to a higher‑quality data mix and improved token efficiency. Key innovations include Sparse Gated Attention for efficient long‑context handling, Gated Norm for training stability, and a Think‑Fusion recipe that allows switching between thinking and non‑thinking modes within the same model.
By Cheolseung Baek, Dhammiko Arya, Eunki Kim, Gun Song, Gyoungeun Han, Hyunho Yang, Hyunjun Eun, Jin Kim, Junyoung Park, Juyun Wee, Minki Hong, Minkyung Park, Minsang Kim, Minsoo Kang, SaeRom Kim, Sangjin Kim, Sangyeol Lee, Seojin Lee, Seokhwan Jo, Seokyoung Hong, Seongho Choi, Seonghye Cho, Seongmin Ok, Sereimony Sek, Seungmo Cho, Seungsik Kim, Singon Kim, Sohee Park, Sooyeon Park, Subin Yi, Sungbin Yoon, Sungeun Lee, Sung Jun Cheon, Sungwan Kim, Sunwoo Lee, Tae Yoon Kim, Wonbeom Jang, Yohan Ra, Yong-jin Han, Youngjin Kim, Youngrang Kim, Yujin Kang, Yujin Lee
arXiv:2508. 16771v3 Announce Type: replace-cross Abstract: Code Language Models (CodeLLMs) learn token importance from data correlations, whereas human developers attend selectively to semantically salient code.
By Yifan Zhang, Chen Huang, Yueke Zhang, Jiahao Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang
arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?
By Ely Hahami, Ishaan Sinha, Lavik Jain
arXiv:2606. 16494v1 Announce Type: cross Abstract: Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by conditioning a reader on passages retrieved from a Wikipedia-scale knowledge base.
By Jieyuan Liu, Jianyang Gu, Shijie Chen, Jefferson Chen, Zhen Wang
arXiv:2609.14207v1 Announce Type: new
Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental...
By T\'ea Wright, Alane Suhr
arXiv:2609.00746v1 Announce Type: new
Abstract: Fine-tuning a pretrained LLM into a vision-language model (VLM) can erode the backbone's text capability, with the damage concentrated on tasks that re...
By Minsik Choi, Geewook Kim, Young Geun Kim
arXiv:2608. 11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways.
By Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon