MaskCoFT introduces a masked co‑adaptive fine‑tuning approach for mixture‑of‑experts language models, training both routers and experts jointly with a cross‑entropy loss. A learnable binary mask limits each layer’s Top‑K routing to a subset of experts during fine‑tuning, allowing experts to adapt to the tokens they receive. In simulated GPU cache scenarios, MaskCoFT reduces expert fetches per token by 23.7% for Mixtral‑8x7B and 10.1% for DeepSeek‑V2‑Lite, and lowers inference time per output token by up to 16.4% and 5.5% respectively, while maintaining or improving accuracy across nine benchmarks.
By Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
The paper argues that the causality of language models may be unnecessary or suboptimal when system behavior—extra dominant factors beyond data distribution—is treated as a first‑principle Bayesian feature. It introduces the SBD framework, incorporating system behavior into the evidence lower bound, and demonstrates a counter‑intuitive Causality Tax where ignoring these factors leads to structural error. Using a non‑causal variational family called Green Shell, the authors show through theoretical bounds, implicit measurements, and Neural Tangent Kernel analysis that this approach yields tighter error bounds and improved generalization compared to causal models.
By Xianzhi Zeng, Jiangneng Li, Gao Cong
The article discusses how large language models can now generate user interfaces from natural‑language descriptions, and proposes a framework that turns HCI design knowledge into machine‑readable ‘skills’ that guide on‑demand UI construction. By treating the dialogue between user and AI as a shared cognitive workspace, these skills encode established HCI principles (e.g., Nielsen’s heuristics, WCAG criteria) as editable, version‑controlled artifacts. The authors outline a research agenda that envisions HCI moving from heuristic checklists to a machine‑executable, open‑source craft.
By Nathan Conklin, Miranda Capra, Chris North
The paper introduces VERSE, a Verified Self‑Evolving optimizer that enhances LLM agent harnesses by allowing the optimizer to test edits, replay failures, and perturb steps while tracking fixes and regressions. VERSE builds its own tools for failure analysis, verification, training audits, and workflow control, and uses this feedback to revise the harness’s prompts, skills, tools, hooks, and notes without changing model weights. In experiments across five executors and multiple languages, VERSE improves all evaluated harness optimizers, achieving higher accuracy on held‑out and out‑of‑distribution tasks compared to the strongest baselines.
By Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy
The paper examines how transformer policies, which process agents as ordered token sequences, perform in multi‑agent robot learning where agent teams are unordered. It finds that low permutation error can mask action collapse, where all agents choose the same action, and proposes additional diagnostics such as action diversity and same‑action fraction. Experiments show that while a PPO‑ID baseline avoids collapse, it remains order‑sensitive, and that strong equivariance regularization can still cause homogeneous behavior; a weak penalty improves robustness and preserves diversity for three‑agent teams, but four‑agent teams need much smaller regularization weights.
By Amit Thakur, Mukesh Singhal
The paper introduces SymCE, a dataset of 4,707 false undergraduate‑algebra and real‑analysis conjectures each paired with a deterministic Python verifier that can be executed to check truth. It studies counterexample generation by training a 4B‑parameter language model (Qwen3‑4B) first with supervised fine‑tuning and then with reinforcement learning (GRPO) using the verifier as a reward. The results show that counterexample‑only fine‑tuning can destroy true‑theorem recognition, while reinforcement learning with a sparse outcome‑only reward restores and surpasses baseline performance, achieving higher success rates than larger open‑weight models and competitive performance on other math benchmarks.
"whyItMatters":"The study demonstrates that reinforcement learning guided by a per‑theorem verifier can effectively repair the shortcomings of supervised fine‑tuning in theorem proving, offering a practical approach to improve language models’ mathematical reasoning capabilities."
By Omar Farouk Zouak, Houssam Eddine Boukhalfa, Soumaya Lakehal, Shiv Katiyar, Samia Nefti-Meziani
The paper investigates whether clinical concepts are represented as distinct, causally influential centers within the latent space of large language models (LLMs). By evaluating eleven open-weight LLMs, the authors discover that each model contains dedicated clinical concept centers that are interpretable, activate only on relevant clinical narratives, and drive model behavior in both constrained and open-ended contexts. These centers can be leveraged for evaluation and performance improvement, as steering models along them enhances downstream clinical outcomes and aligns with clinician preferences.
By Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler
The paper introduces SBERT2S1, a method that transforms Sentence-Transformers encoders into typed decision models for biomedical text, and presents BIODECIDE, a suite for evaluating such models, along with MEDLINE‑S1, a large training set of 243k decisions derived from NLM indexing. Experiments show that retrieval‑trained encoders improve zero‑shot matching of content‑bearing options, and that a prior‑fused residual (PFR) head benefits most from retrieval pre‑training, while a cross‑head (C) head generally outperforms PFR across all objectives. The authors also release code, data, and a model, and discuss calibration and reward‑normalisation effects in their RLCD recipe.
By Pritam Deka
APDMem is a hierarchical long‑term memory system for personalized LLM assistants that uses progressive disclosure to retrieve conversation history. It organizes memory into four layers—thematic summaries, personalized key facts, turn‑level evidence notes, and raw messages—and a controller reads from the top layer, drilling down only when necessary. This approach balances cost and fidelity, allowing simple queries to finish early while deeper inspection is triggered for complex or exact‑evidence requests, and a note synthesizer structures retrieved evidence before answer generation. Experiments on LongMemEval show APDMem performs strongly while accessing only 8% of the total conversations.
By Chin-Lun Fu, Anagha Kulkarni, Hong Ni, Behrouz Madahian
CuBEs introduces culturally‑situated behavioral evaluations for large language models, adding cultural context to test scenarios and assessments. The authors built a human‑labeled dataset covering 12 cultures, revealing significant cross‑cultural differences that one‑size‑fits‑all judgments miss. Evaluating 13 LLMs shows that culturally situated tests uncover varied behaviors, such as Western political bias versus non‑Western religious or colonial biases, which standard evaluations overlook.
By Hoda Ayad, Tanu Mitra, Abhishek Mukherji
XGenAct is a world action model that encodes RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos using deterministic codecs. It trains a single video diffusion transformer with a unified objective, sampling different perception and action tasks to learn temporal prediction across multiple spatial modalities without separate heads. Experiments on RLBench show that structured perception training boosts closed‑loop success, with XGenAct achieving 52% success on five external tasks compared to 26% for the best baselines, and it outperforms pipelines that first generate RGB and then apply a frozen perception expert for depth and segmentation prediction.
By Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li
The paper introduces a fully automatic pipeline for 3D dendrite instance segmentation in serial block-face scanning electron microscopy (SBF-SEM). It combines YOLOv6-guided Segment Anything Model prompting, iterative 2D mask refinement, random forest 3D linking, and high‑resolution nnU‑Net refinement to produce coherent, well‑separated dendrite reconstructions without manual prompting. Applied to hippocampal CA1 data from a control and an epileptic rat, the method achieves high semantic accuracy (Dice 0.93/0.91) and strong instance performance, while highlighting dense‑region recognition as the main challenge in epileptic tissue.
By Zewen Zhuo, Ilya Belevich, Eija Jokitalo, Alejandra Sierra, Jussi Tohka
ProgressNet is a training‑free framework that enables a frozen text‑to‑image model to follow a sketching session in real time. It allows strokes to be added or erased, prompts to be revised, and produces an updated image in about a second per turn without adding new parameters. The method uses three inference‑time mechanisms—Previous‑Concept Memory, Layer‑Selective K/V Injection, and Banded Adaptive Control—to maintain fidelity and coherence across sketch domains, outperforming existing approaches especially when sketches are partially erased.
By Arkaprabha Basu, Chaitat Utintu, Yi-Zhe Song
HakushoBench is a Japanese chart and table visual question answering benchmark created from 33 governmental white papers, comprising 2,053 images across more than ten types. The dataset includes manually annotated and independently verified QA pairs that evaluate holistic understanding of charts and tables rather than just local visual cues. Experiments show that HakushoBench is significantly harder than existing Japanese benchmarks, with sub‑10B open‑weight models achieving at most 58.6% accuracy and even large models like Qwen3.5-397B-A17B lagging behind Gemini‑3‑Pro by 8.1 points.
By Issa Sugiura, Shuhei Kurita, Yusuke Oda, Naoaki Okazaki
Large language models (LLMs) are increasingly used for clinical reasoning, yet their ability to revise judgments as patient evidence changes is uncertain. In this study, researchers evaluated longitudinal belief updating using real intensive‑care patient trajectories and found that conditioning on prior judgments often increased prediction error. Controlled experiments revealed two failure modes: models were more responsive to worsening than improving respiratory evidence, and they shifted estimates significantly when prior risk was altered, indicating a causal influence of prior beliefs. Prompting did not improve reliability, and an Evidence‑Validated Longitudinal Update (EVLU) approach produced fewer but more trustworthy revisions, highlighting a reliability‑coverage trade‑off.
By Min Zeng, Rui Zhang
The paper studies how language‑model agents adopt beliefs, measuring the probability that an agent accepts a claim based on the number of peers endorsing it. The adoption curve is sigmoid, indicating complex contagion, with a threshold influenced by the claim’s plausibility, the source’s reliability, and the agent’s disposition—dimensions that can be collapsed into a single coherence metric relative to the agent’s prior beliefs. In networks of AI agents, belief spread is stronger on clustered than random networks, and the system shows a bifurcating cascade window and hysteretic consensus that makes consensus difficult to reverse once formed.
By Tathagata Banerjee, Nima Moghaddas
The paper introduces MECo, a large language model–driven multi‑task evolutionary framework that enables zero‑shot cross‑problem generalization for combinatorial optimization. MECo maintains task‑conditioned heuristic populations, uses a transfer gap based on cross‑task performance to guide interactions, and selects a compact heuristic set that covers source combinations. Experiments on 32 variants of vehicle routing and flexible job‑shop scheduling demonstrate that MECo achieves lower mean costs than eight automated heuristic design baselines and improves both in‑domain and out‑of‑domain performance when integrated with other methods.
By Zhouliang Xie, Changliang Zhou, Genghui Li, Zhenkun Wang
CVE2AP is an LLM-based system that automatically converts natural language CVE descriptions into PDDL-encoded attack paths. It uses structured prompting and an error‑feedback loop that refines outputs based on planner‑reported syntactic and solvability errors. Empirical tests across various LLMs show high quality results, with up to 86.9% syntax correctness, 78.6% solvability, and 93.1% semantic correctness, and GPT‑5.5 providing the best quality‑cost balance.
By Lin Cui, Vincenzo Scotti, Raffaela Mirandola
The paper investigates how training large language models to use fewer tokens in chain-of-thought (CoT) reasoning impacts the faithfulness and monitorability of the generated explanations. Three efficiency methods—fixed generation budget, per-example length target, and group-relative length reward—were applied during fine‑tuning, and the resulting models were evaluated on how well their CoT reflects decision processes and whether it signals changes due to input interventions. Results show that while faithfulness generally decreases because models become less consistent, monitorability remains relatively robust, with models still indicating the influence of input changes even when CoT is shortened.
By Samuel Lewis-Lim, Xingwei Tan, Mario Sanger, Zhixue Zhao, Nikolaos Aletras
The paper introduces CERTID, a formal identification pipeline that uses the sound and complete causal identification algorithm ID to certify whether a causal effect is identifiable from a given graph and query. CERTID also verifies returned formulas against structural causal models with known interventional distributions, mitigating structural leakage and providing grading guarantees. The authors evaluate three frontier reasoning models on 1,200 certified instances, finding that accuracy is a poor proxy for soundness, with false-claim rates on non-identifiable queries varying by seventeen-fold across models.
By Arman Behnam, Binghui Wang