The paper introduces CAST, a concept-guided artifact suppression tuning framework that uses sparse autoencoders to identify and suppress note-specific artifacts in clinical language models. CAST labels latent features with an LLM-assisted pipeline and ICD‑10 constraints, then fine‑tunes the model while providing post‑hoc per‑concept attributions for auditability. In experiments on MIMIC‑IV discharge‑note mortality prediction, CAST outperforms standard fine‑tuned encoders and competes with strong LLM baselines while offering a feature‑level audit trail of clinical concepts and suppressed artifacts.
By Jin Mu, Guanhua Chen
The study investigates how document segmentation (chunking) and embedding choices influence Retrieval-Augmented Generation (RAG) performance for Turkish, a morphologically rich language. Using a fully crossed experimental design, the authors compare three chunking strategies (fixed-length, semantic, and layout-aware Docling), five embedding models, and two large language model generators across three documents with different layouts, generating 9,000 graded question-answer evaluations. Key findings include that layout-aware chunking reduces the impact of embedding choice, the top embedding models perform similarly, faster generators are not more accurate, and the optimal configuration varies with content type, achieving a best overall accuracy of 87.0%.
By Mustafa Serta\c{c} T\"urkel, Fatma Nur Korkmaz, Ahmet Tu\u{g}rul Bayrak
The paper investigates how the language of prompts and responses affects large language model (LLM) outputs. Using five models and 68 non‑translation questions, the authors compare English‑to‑English, English‑to‑Norwegian, Norwegian‑to‑Norwegian, and Norwegian‑to‑English conditions, yielding 1,348 responses after filtering. They find that prompt language strongly influences response length—Norwegian prompts shorten English outputs by ~37 % and English prompts shorten Norwegian outputs by ~41 %—while semantic similarity remains high across conditions.
By Thi Thanh Nhan Nguyen, Mai Khoi Tieu, Michael A. Riegler, P{\aa}l Halvorsen, Thu Nguyen
The study evaluates how extractive prompt compressors affect token costs across ten languages, finding that compressors trained on English data widen the token premium gap for non‑English languages, while a multilingual compressor does not. The gap is tied to the supervision data rather than model architecture, and aggressive compression can reduce non‑English contexts to near‑zero utility. A translate‑then‑compress approach can match or outperform native compression at roughly half the token cost in several languages.
By Mantas Lukauskas
The paper proposes a fragment‑based reasoning framework for large language model–based machine translation. It extracts parallel source‑target fragments from retrieved similar examples and uses these fragments as intermediate reasoning traces to generate the final translation. Experiments with the Qwen3 model across six languages and multiple domains show that this approach outperforms standard k‑shot or basic drafting methods.
By Maxime Bouthors, Josep Crego, Fran\c{c}ois Yvon
The paper introduces a modular data‑science pipeline that estimates public sentiment toward individuals using fragmented, unstructured open‑source intelligence. The pipeline combines web search, text extraction, relevance filtering, tokenisation, co‑reference resolution, and sentiment analysis to produce auditable person‑level sentiment distributions. By comparing AFINN, VADER, and the domain‑specific MINOS algorithm, the authors show that MINOS best distinguishes positive, ambiguous, and negative reputational cases, and they apply the method to the UK Honours system to support transparent, reproducible, human‑in‑the‑loop sentiment assessment for high‑stakes decisions.
By Francesca von Braun-Bates, Sunreeta Sen, Indraayudh Talukdar, Anirban Lahiri
RuleWeaver is a benchmark construction framework designed to evaluate large language models’ ability to reason over complex, rule‑centered scenarios. It begins with corpus‑derived IF‑THEN meta rules, expands them into more intricate rules, and composes these into scenario‑based QA instances. The benchmark assesses not only final answer correctness but also process‑level metrics such as rubric‑based answer quality, rule recall, and rule precision, revealing that current LLMs achieve only about 50% of the maximum rubric score on these tasks.
By Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu
Cascaded Batch Prompting introduces a two‑stage method that separates complex reasoning from symbol grounding to address the unpredictability of conventional batch prompting. Experiments on multiple‑choice question answering and natural language inference show that this approach outperforms standard single prompting while maintaining a speedup proportional to batch size. The technique establishes a new state‑of‑the‑art position on the Pareto frontier for efficiency and performance.
By Sho Hoshino, Peinan Zhang
The paper introduces MAPLE, a family of decoder‑only language models pretrained with document‑level geographic metadata such as source URL, country, and continent. MAPLE is evaluated on a new benchmark, LocalNewsQA, which tests whether models can switch answers when the locale changes. Experiments show that, with inference‑time metadata fixed, MAPLE outperforms metadata‑free controls in both answer switching and accuracy on locale‑dependent questions, and these gains grow with model size.
By Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav
HALO is a heterogeneity‑aware, language‑aligned foundation model for inertial measurement unit (IMU) based human activity recognition. It uses a two‑stage training process: first, a self‑supervised encoder learns to handle diverse sensor configurations and natural‑language sensor descriptions; second, the encoder is aligned with text embeddings through synonym‑aware contrastive learning, enabling open‑set recognition via cosine similarity. Trained on ten public HAR datasets, HALO outperforms five state‑of‑the‑art baselines across eight metrics while using only ~35 M parameters, and improves zero‑shot open‑set accuracy by 13.7 percentage points over 87 training labels.
By Zihan Ding, Liyu Zhang, Xiaomin Ouyang
The paper investigates federated adversarial training (AT) for vision transformers, a topic not previously explored in federated learning (FL). It evaluates various transformer architectures and aggregation strategies, and introduces FedWAvg, an extension of FedAvg that weights client updates based on similarity of their last-layer representations. Experiments demonstrate that FedWAvg yields higher robust accuracy than existing aggregation methods in non‑IID settings.
By Ahmed Aldahdooh, Wassim Hamidouche, Olivier D\'eforges
The paper introduces PACE, a factor-guided, progressive framework for acquiring evidence in long-video question answering. PACE first indexes clip-level descriptions using question-derived factors, then refines evidence retrieval with contrastive cues derived from candidate answers. On the MMR‑V dataset, PACE achieves 42.6% accuracy and recovers 66.9% of annotated cues, outperforming direct inference and prior agentic baselines, and shows consistent improvements across several long-video benchmarks.
By Baixuan Xu, Yinyui Xu, Tianshi Zheng, Zhaowei Wang, Weiqi Wang, Haochen Shi, Jiayu Liu, Qing Zong, Xiyu Ren, Xinyu Geng, Zhitao He, Yangqiu Song
Cascaded Batch Prompting introduces a two‑stage method to improve large language model inference by separating complex reasoning from symbol grounding, addressing the unpredictability of traditional batch prompting. Experiments on multiple‑choice question answering and natural language inference show that this approach outperforms single prompting while scaling speed with batch size, achieving a new state‑of‑the‑art balance between accuracy and efficiency.
TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a family of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can remove up to 96% of training tokens without harming quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this mixture outperform much larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.
By Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir
FLEET is a token‑based feature extractor that processes event camera data directly, using random Fourier features and cross‑attention to compress variable‑length event streams into fixed‑size latent representations. By decoupling inference cost from sensor resolution, it avoids the high compute and temporal blurring associated with CNN‑based grid aggregation. Experiments on a new high‑throughput benchmark show that FLEET outperforms state‑of‑the‑art methods and remains robust across different observation frequencies.
By Tristan Gottwald, Maximilian Schier, Melanie Schaller, Bodo Rosenhahn
The paper reports the first use of simulation‑based inference for resonant inelastic X‑ray scattering (RIXS) spectroscopy, applying truncated marginal neural ratio estimation and conditional flow matching to infer full posterior distributions of Hamiltonian parameters for two Ni$^{2+}$ compounds. A vision‑transformer encoder tailored to the RIXS map’s physical layout produces sharper, better‑covered posteriors than generic image encoders. The validated method, applied to experimental data, uncovers parameter correlations invisible to point estimators and yields posterior predictive distributions that closely match observed spectra, enabling new analyses such as nuisance‑marginalized uncertainty quantification, multi‑measurement posterior fusion, and active experimental design.
By Samuel Klein, Thomas M. Linker, Louis Conreux, Daniel Ratner, Apurva Mehta, Makoto Tachibana, Jiemin Li, Jonathan Pelliciari, Valentina Bisogni, Wei He, Xiangpeng Luo, Mark P. M. Dean, Marton K. Lajer, Michael Kagan, Joshua J. Turner, Yongqiang Cheng, Sean Gasiorowski
The paper presents an instruction‑tuned large language model (LLM) based on Qwen3 that is fine‑tuned for hate speech mitigation by unifying 36 English hate speech datasets. The authors show that this generalist LLM achieves state‑of‑the‑art performance on in‑domain benchmarks and delivers significant gains in cross‑domain and cross‑lingual generalization, outperforming specialist encoder‑based classifiers.
By Lukas Edman, Daryna Dementieva, Alexander Fraser
The paper "Demystifying Reinforcement Learning Post-Training of Language Models" investigates how reinforcement learning (RL) post‑training enhances large language models (LLMs) for tasks such as reasoning, math, and coding. By isolating RL components in a controlled setting, the authors analyze how the base model’s prior distribution, reward granularity, prompt diversity, and model scale influence outcomes, using policy entropy to compare pre‑training, supervised fine‑tuning (SFT), and RL stages. The study clarifies the role of spurious rewards, the importance of the base model’s probability mass on desired behaviors, and how these factors interact to determine post‑training success, offering a practical primer for NLP researchers.
"whyItMatters":"The work provides a clearer understanding of RL post‑training mechanics, helping researchers and practitioners effectively apply RL to improve LLM capabilities."
By Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh, Natasha Jaques
Loss-Based Active Learning for Neural Abstractive Summarization proposes LOBSTER, an active learning framework that selects unlabeled documents similar to the model’s high‑loss training examples to correct specific weaknesses. The method is tailored for abstractive summarization, addressing instability and computational bottlenecks seen in prior work. Experiments on three benchmark datasets and two backbone models show that LOBSTER matches or surpasses state‑of‑the‑art performance while speeding up query selection by up to 665×.
By Michail Ioannou, Tatiana Passali, George Michalopoulos, Grigorios Tsoumakas