The study used causal discovery and mediation analysis on 42,293 UK Biobank participants to map how moderate‑to‑vigorous physical activity (MVPA) reduces dementia risk. It found interconnected pathways involving depression, functional capacity, smoking, hypertension, cardiovascular disease, chronic kidney disease, and brain injury, with depression accounting for 15.1% of the effect. Sex‑specific analyses showed females mainly linked to metabolic and functional pathways, while males were more influenced by behavioral and cardiometabolic cascades.
By Wasif Khan, Panayiotis V. Benos, Joshua K. Wong, Ruogu Fang
The paper introduces the Hop-Decayed Influence (HDI) attack, which targets auxiliary schema-level structures—semantic summaries, hierarchical edges, and pre-computed scores—used in GraphRAG pipelines for retrieval prioritisation. By propagating query-aware influence, HDI identifies high-impact targets and corrupts a minuscule fraction (as low as 0.016%) of these structures, achieving an 88–94% success rate across HotpotQA and 2WikiMultiHopQA benchmarks on Microsoft GraphRAG and HippoRAG2 architectures. The modifications affect up to six queries each, yielding a 1:N amplification that instance-level attacks cannot achieve, and they evade perplexity and paraphrase defenses with over 99% evasion, exposing a structural blind spot in current GraphRAG defenses.
By Jisung Park, John Le, Heath Cooper
The paper introduces Inner Momentum (IM) for differentially private Muon training. It addresses distortion caused by per-example gradient clipping by averaging each example’s Muon gradient over the current and recent models before clipping, thereby bounding clipping-induced distortion. Experiments on private GPT‑2 fine‑tuning show that DP‑Muon‑IM consistently improves BLEU and ROUGE‑L scores and reduces polar error compared to standard DP‑Muon.
By Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai
EVOL is a framework that uses a knowledge‑tracing simulator to generate per‑learner expert demonstrations via evolutionary search, and then trains a deployment‑free policy that distills these demonstrations into a feed‑forward learner. The method employs an asymmetric actor‑critic architecture, where the actor plans blindly as in real deployment while the critic has access to privileged simulator state during training. Across three educational datasets and varying path lengths, EVOL outperforms eight baselines and shows that the quality of evolutionary experts, rather than the imitation objective, drives final performance.
By Geonwoo Bang, Dongho Kim, Moohong Min
The paper introduces Prefix‑Certificate Retention (PCR), an exact optimization framework for deciding which prefix states of a language model to cache in order to balance recomputation and storage time. PCR models usable‑prefix rewards as nodes with prerequisites tied to timeout thresholds and preceding hit certificates, reducing the problem to a single minimum‑cut graph whose size scales linearly with block lookups and timeout choices. The authors provide a breakpoint theorem that extends the construction to all nonnegative timeouts without discretization error, a linear‑time dynamic program for ordered timeouts, and empirical validation on 39,632 public Mooncake requests, showing that ordered timeouts achieve the unrestricted optimum in most trace‑grouping cases and that heterogeneous retention can improve memory‑time tradeoffs.
"whyItMatters":"The study offers a tractable, auditable optimization model for retention policies that directly measures usable prefix blocks and storage time, providing a concrete benchmark for evaluating memory‑time tradeoffs in language‑model serving."
By Shivam Gupta
The paper investigates how to adapt large language models to mimic an individual author's style using only a few example abstracts, a task made harder by the formal nature of scientific writing. It proposes three style‑conditioning methods—contrastive activation steering, a network predicting steering vectors, and a hypernetwork predicting LoRA adapters—and finds that while fine‑tuning captures the strongest style signal, it harms fluency; the hypernetwork offers the best balance between style imitation and output quality for both seen and unseen authors. The authors also show that steering can be performed at the author level by contrasting author abstracts against style‑neutral generations for the same content, eliminating the need for a predefined style inventory and outperforming inventory‑based approaches.
whyItMatters":"The study provides practical techniques for author‑style transfer in scientific writing, revealing a trade‑off between style fidelity and fluency and demonstrating that hypernetworks can effectively balance these aspects."
By Leonard Popp, Danni Liu, Supriti Sinhamahapatra, Jan Niehues
Asterism is a tool that extracts observations from hundreds of papers as concept‑relation triples and unifies these concepts in a hierarchical ontology. Researchers can curate an evidence graph, aggregate observations at various levels of granularity, and focus theory formation on phenomena that match their preferences. In a field deployment with ten researchers, and in two case studies involving immunology and agriculture, teams used Asterism to discover new mechanisms and generate hypotheses for follow‑up experiments.
By Joseph Chee Chang, Michael D'Arcy, Amy X. Zhang, Pao Siangliulue, Sangho Suh, Aakanksha Naik, Jena D. Hwang, Javier Ramos Benitez, Stella Wroblewski, Matt Latzke, Michael Cuoco, Ruben Lozano-Aguilera, Kris Ganjam, Joel Chan, Doug Downey, Peter Jansen, Kyle J. Travaglini, Daniel S. Weld
The paper presents a system for the Touché 2026 causality extraction challenge, focusing on counter‑causal claims—sentences that appear causal but actually deny causation. It tackles three subtasks: detecting causal sentences, extracting cause and effect spans, and labeling polarity (procausal, counter‑causal, or uncausal). The approach uses a fine‑tuned classifier with a cross‑task rule for detection, an ensemble of three RoBERTa‑large BILOU+CRF taggers for extraction, and counter‑causal data augmentation via a large language model for polarity classification, achieving state‑of‑the‑art scores on the CCNC test set.
By Roham Zendehdel Nobari, Shayan Sooratgar
The paper presents a system for matching job candidates to vacancies that provides interpretable evidence rather than a single relevance score. It uses a two‑stage approach: an LLM‑based labeler refined through recruiter feedback and a distilled bi‑encoder that runs online on CPU. The model, trained on 168,772 labeled pairs, achieves 95.79% agreement with recruiter‑recorded decisions on a production‑feedback subset.
By Ilya Chekin (BroutonLab), Vyacheslav Malyugin (BroutonLab), Vladimir Chirkov (BroutonLab), Mikhail Yurushkin (Curately)
The paper introduces a generative‑informed neuro‑symbolic framework that combines generative syntactic theory with AraBERT to resolve structural ambiguity in Modern Standard Arabic noun phrases. By treating ambiguity as a candidate‑based decision task, the model explicitly constructs and evaluates linguistically motivated alternatives, achieving high accuracy (96.88%) and strong F1 scores on an unseen evaluation set. Analysis shows uneven performance across attachment types, with near‑perfect recall for high/VP attachment but lower recall for low/NP/embedded attachment, highlighting challenges in recovering embedded interpretations.
By Mohammed Damom, Muneef Y. Alshawsh, Ashraf A. Naji, Mustafa Ali Alhamzi, Fawwaz An-Nashef, Jameel Ahmed Elayah, Mohammed Q. Shormani, Noman AL-Sayadi
HakemBench is a Turkish benchmark for typed decision-making tasks, comprising 2,346 items and 4,275 choice, yes/no, and score questions across seven tracks (fact‑check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). The benchmark evaluates models on decision quality (macro F1), calibration (normalised Brier score), and selective automation (normalised area under the generalised risk‑coverage curve), combining these metrics by a geometric mean and reporting confidence intervals from 2,000 bootstrap draws. Gold labels are generated from blind passes of a single AI model family compared with votes from other large language model families, and the benchmark’s leader achieved a composite score of 0.888 while the lab’s own model ranked 7th with 0.660.
By Sait Furkan Teke (ufak AI)
FiberGeoText (FGT) is a vision‑language model that clusters short‑range superficial white matter streamlines from ultra‑high‑resolution diffusion MRI into population‑level groups. It jointly encodes each streamline’s 3‑D trajectory, cortical anatomical context (via text from multiple parcellation schemes), and shape, using a pretrained large language model to unify heterogeneous anatomical descriptions. Evaluations on 0.76 mm diffusion data show that FGT outperforms state‑of‑the‑art methods in cortical parcel coherence, shape consistency, cluster‑size consistency, and cross‑subject correspondence, and it generalizes well to unseen subjects, recovering 96.7 % of learned clusters.
By Yuqian Chen, R. Jarrett Rushmore, Guikun Chen, Fan Zhang, Edward Yeterian, Nikos Makris, Yogesh Rathi, Lauren J. O'Donnell
The paper introduces AURA, an LLM-powered mask‑reconstruct framework for anonymizing text while preserving utility. It decouples privacy localization from utility‑preserving reconstruction and uses adversarial checks to select candidates. Experiments on real‑user interview transcripts show that AURA achieves the lowest agentic re‑identification rates among non‑DP methods and retains more contextual utility than prior LLM anonymizers.
By Ziwen Li, Jianing Wen, Tianshi Li
ZAGNet is a Zone‑Aware Graph Neural Network that models temporally tracked lung ultrasound findings as graph nodes linked by anatomical zone adjacency, enabling contextual propagation via a graph transformer and a virtual global node for patient‑level diagnosis. It handles missing zones by operating on a flexible graph structure, and outperforms traditional max/mean pooling on a multicenter dataset of 714 subjects, achieving AUCs of 0.803 for consolidation and 0.893 for pleural effusion. The study demonstrates that graph‑based inter‑zone reasoning improves automated patient‑level LUS assessment.
By Li Chen, Shubham Patil, Rashid Al Mukaddim, Jochen Kruecker, Balasundar Raju, Alvin Chen
The paper introduces iS-KV, an online low‑rank KV‑cache compression technique that uses block‑incremental SVD to manage memory during long‑horizon autoregressive decoding. Unlike token‑eviction methods, iS‑KV retains all positions in a compact representation by keeping a recent window exact and incrementally folding older states into bounded‑rank bases, synchronizing coordinates as the basis evolves. Experiments on DeepSeek‑R1‑Distill‑Llama‑8B and Qwen3‑8B show that iS‑KV achieves high accuracy (82.6% and 89.2% respectively) while providing 4.06‑fold and 5.64‑fold compression, outperforming token‑eviction baselines under matched memory budgets.
By Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
The paper introduces Test-Time Calibration Learning (TTCL), a label‑free framework that adapts both reasoning accuracy and verbalized confidence of large language models directly on unlabeled target‑task data. TTCL generates self‑supervision signals from multiple model responses, enabling calibration without ground‑truth labels and proving theoretically as a bounded surrogate for the ideal calibration objective. Experiments on mathematical reasoning and factual question answering show consistent improvements, with base models gaining an average 40.13% in accuracy and 70.80% in ECE reduction across eight benchmarks, and further gains under domain shift.
By Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han
The paper introduces FOVEATED, a plug‑and‑play framework that improves atomic‑fact recall in unstructured knowledge editing (UKE) for large language models. By randomly shifting Rotary Position Embedding (RoPE) positions during editing, FOVEATED creates focused views of each sentence, counteracting the context‑reliance problem where edited LLMs reproduce passages but fail to recall individual facts. Experiments show consistent gains across five editors, two LLM backbones, and three benchmarks.
By Ding Wu, Ye Zhang, Haoyu Wang, Tianci Liu
The paper introduces Causal-Invariant Masking (CIM) to better quantify epistemic uncertainty in Multimodal Large Language Models (MLLMs) by measuring semantic shift between original predictions and those conditioned on a causally-focused view. It proposes Semantic Divergence as a core metric that converges to the variance of the model’s sensitivity to non‑causal correlations, and introduces Expected Embedding Drift (EED) as a fast geometric proxy that estimates this shift directly in the embedding space. Experiments demonstrate state‑of‑the‑art uncertainty quantification performance and a nearly 50% speedup with EED.
By Haoyang Luo, Linwei Tao, Jie Gui, Xinghao Chen, Chang Xu, Jianyuan Guo, Minjing Dong
Enrich-on-Graph (EoG) is a flexible framework that uses large language models to enrich knowledge graphs, thereby bridging the semantic gap between structured graphs and unstructured queries in complex reasoning tasks. By leveraging LLMs’ prior knowledge, EoG enables efficient evidence extraction from knowledge graphs, achieving precise and robust reasoning while maintaining low computational costs and scalability. The authors also introduce three graph quality evaluation metrics for query‑graph alignment, theoretically validate their optimization objectives, and demonstrate state‑of‑the‑art performance on two KGQA benchmark datasets.
By Songze Li, Zhiqiang Liu, Zhengke Gui, Huajun Chen, Wen Zhang
LiteEMG-FM is an efficient hybrid CNN‑Transformer foundation model designed for electromyography (EMG) sensing. It is pretrained on 16 diverse upper‑ and lower‑limb EMG datasets, enabling representations that generalize across users and datasets. The model incorporates a lightweight, always‑on 1D‑CNN wake‑up module that filters rest and non‑target activity, activating LiteEMG‑FM only for valid gestures, and is evaluated in full inference offloading, split inference, and full on‑device processing scenarios, showing superior performance over state‑of‑the‑art time‑series foundation models and supervised baselines, especially in zero‑calibration cross‑participant and data‑scarce conditions.
By Tianhao Wu, Xu Wu, Amirmohammad Radmehr, Jiawei Yu, Yi Wu, Phuc Nguyen, Jian Liu