Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.
By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
The paper investigates whether semantic entropy—a measure of disagreement among a language model’s sampled answers—can serve as a cheap signal for deciding when to route a query from a small to a larger language model. Experiments on GSM8K and other benchmarks show that semantic entropy can distinguish small‑model mistakes and improve routed accuracy, but the authors also reveal that a simple question‑difficulty rule can mimic its performance and that other factors (definition of success, benchmark design, live sampling cost) can undermine its effectiveness. They propose a checklist of checks to validate escalation signals and demonstrate how to predict when a cached‑outcome approach will fail.
"whyItMatters":"The study highlights the importance of rigorous evaluation of escalation signals, showing that seemingly promising metrics can be misleading without proper controls and that practical routing decisions must account for cost and benchmark design."
By Ramin Pishehvar, Andrea Morandi, Mahesh Viswanathan
The paper presents a defense‑in‑depth framework for large language models that separates internal activation steering from external memory handling to combat sycophancy. It evaluates four open‑weight models on a new MemSyco‑Bench dataset, testing five memory‑defense configurations—including a Router Gate that selectively rewrites, keeps, or drops memories—and measures sycophancy and accuracy across 1,550 items. Results show that selective Router Gate filtering preserves more accuracy than complete memory removal, while inverse steering slightly reduces sycophancy but is not statistically significant.
By Ritvij Sharma, Russell Dlugosz, Ryan Zhou, Maheep Chaudhary
PsyCIDRA is a dual‑agent framework that couples a free‑form psychiatric interviewer with a diagnostic reasoning agent to support expert review. The interviewer agent uses tools to keep working notes, load expert skills, and pull ICD‑11 references, while the diagnostic agent receives the interview transcript and generates hypotheses with supporting, conflicting, and missing evidence, withholding a final hypothesis if insufficient support exists. In simulations and a blinded human study, PsyCIDRA achieved higher diagnostic agreement and rank‑1 accuracy than direct prompting, indicating its promise for assisting psychiatric assessment through interactive dialogue.
By Milad Mohammadi, Fatemeh Akrami Shamsabadi, Zahra Mohseni, Amirhossein Safdarian, Malekfarhad Malek, Hadi Moradi, Hesham Faili
arXiv:2610.07544v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to judge how well an advertisement matches a query, but the fairness of these judgments has received...
By Weiwei Wang, Yinchuan Xu, Jialu Gao, Youkow Homma, Jian Jiao
arXiv:2610.07570v1 Announce Type: new
Abstract: In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism...
By Xiaoyang Wang, Tianrui Wang, Christopher C. Yang
LOGIC is a benchmark and evaluation framework that tests how language models can ground engineering requests in a deterministic inventory of candidate changes before propagating selected changes through an electrical traceability graph. The benchmark includes 168 scenarios—144 for selection and 24 for abstention—and evaluates three 7–8B models against intent‑agnostic, lexical, and structured‑evidence methods. Results show that structured evidence can achieve perfect candidate F1 on anchored cases, while large language models perform better on relational‑paraphrase cases; however, grounding accuracy drops as candidate inventories grow, and strict evidence gating reduces false positives but may also remove correct selections.
By Muhammad Faraz Shoaib, Muhammad Qasim, Raisulhaq Mohammed Rizwan, Rahmatullah Safdar, Muzammil Adnan Shaik, Abdul Aleem Mohammed
The paper introduces Personal-Agent Mediated Recommendation, a new paradigm where a personal LLM agent uses cross‑platform user history to adjust a platform’s recommendation ranking. It presents MediateRec, a benchmark for evaluating this mediation, and proposes Personal Attribution Mediation Optimization (PAMO) to balance beneficial rescues against harmful overrides. Experiments show that PAMO improves over outcome‑only reinforcement learning, achieving a better rescue‑harm trade‑off on both synthetic and real cross‑platform tests.
By Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley, Xiangjun Fan
The paper introduces LSC-DPO, a variant of Direct Preference Optimization that dynamically controls the learning signal to maintain sensitivity during training. By analyzing the logistic DPO loss geometrically, the authors identify the sigmoid factor as a key learning signal and propose a log‑space framework for stable target‑regime tracking. Experiments on AlpacaEval 2, MT‑Bench, and Anthropic‑HH demonstrate that LSC‑DPO outperforms standard DPO and other preference‑optimization baselines, and a signal‑budget compensation rule further reduces variability across different coefficient initializations.
By Yang Qu, Yusheng Han, Chengjia Feng, Handan Liu
The paper introduces DEFER1, a deterministic-first enforcement system for large‑language‑model based multi‑agent systems that uses 28 checks to block most attacks and refers only a small fraction to human judges. In tests across four domains, DEFER1 reduces attack success from about 30% to roughly 3%, with 78% of attacks blocked deterministically and only a quarter reaching the judges. The study highlights that rules effectively handle clear policy violations while judges address ambiguous intent, but also reveals weaknesses such as a risk‑score gate that misclassifies many proposals.
By Shaswata Mitra, Raj Patel, Subash Neupane, Sudip Mittal, Md Rayhanur Rahman, Shahram Rahimi
The paper identifies a single input embedding channel, called the massive activation gating channel (MAGC), that controls the emergence of massive activations in large language models. When the MAGC value is sufficiently large or small, the spike feed‑forward network outputs exhibit exceptionally large magnitudes. The authors verify MAGC across six models and provide a theoretical explanation linking the channel to a quadratic form that mixes specific columns of the down‑projection matrix, which produce massive activations.
By Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang
ST-Bench is a new benchmark that tests whether multi‑agent systems (MAS) outperform single agents on complex scientific data analysis tasks. It includes 100 Earth‑science data‑science tasks expanded into 2,067 queries, validated by domain experts. Evaluations show that most MAS configurations beat the cheapest single‑agent baseline, with the best reaching nearly three times its score, though at higher inference cost.
By Qi Cheng, Rongchao Dong, Shengyu Chen, Licheng Liu, Dan Lu, Zhengzhang Chen, Wei Cheng, Yiqun Xie, Haifeng Chen, Xiaowei Jia, Haoyu Wang
OTel is an open telecom AI resource that provides derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, along with 30 full‑parameter post‑trained baselines covering 10 embedding models, 3 rerankers, and 17 language models. The project has seen significant community engagement, with over 16 million model downloads and more than 157 media mentions by May 2026. Post‑training on OTel data improves performance across all model families, achieving 93.1% NDCG@10 for embeddings, 0.947 MRR@10 for rerankers, and 87.8% correctness for language models.
By Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali Maatouk, Leandros Tassiulas, Rex Ying, Nick Sorros, Louis Powell, Nikolaos Vasiloglou, Ashish Vaswani, Somanshu Singla, Adarsh Chaluvaraju
OOPMAS introduces a training‑free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are defined as object‑oriented class definitions with dedicated roles, tools, and persistent state, while workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in‑context improvement without any gradient updates or fine‑tuning, and achieves 89.6% accuracy on a mixed‑task benchmark, outperforming the strongest baseline by 18.1 percentage points.
By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen
The study investigates how large language models (LLMs) compensate for missing financial information by substituting user identity cues. Using 96,600 prompts to Llama‑3.1‑8B‑Instruct, the authors varied the amount of financial facts provided while keeping the underlying finances constant, and measured changes in recommended equity allocations across 100 financial profiles, 138 personas, and seven disclosure conditions. Results show that as financial facts are removed, the influence of identity on advice grows dramatically—from 5 % of variation with full disclosure to 96 % with none—while household size becomes the most reliable predictor when evidence is scarce, and gender effects persist even after controlling for standard errors.
"whyItMatters":"The findings highlight that LLM‑based advisory systems can produce biased financial recommendations when users provide incomplete information, underscoring the need for audits that reflect real‑world disclosure levels and consider the full spectrum of user identities."
By Saanvi Khetan, Sankar Balasubramanian
The paper introduces Agentic Semantic Sensing (Agentic SemS), a closed‑loop framework for AI‑enabled radio access networks that dynamically adjusts sensing configurations within a communication‑feasible profile set. A profile‑conditioned causal Transformer updates semantic beliefs from streaming data, while a semantic utility network guides the selection of the next sensing profile and determines when to exit early, balancing task benefit against sensing cost. Experiments on Widar3.0 demonstrate that Agentic SemS reduces cumulative sensing cost by 25.33% compared to a full‑sequence baseline while maintaining 85.79% Macro‑F1, and that semantic early exit yields an additional 12.35% cost savings with minimal performance loss.
By Zhongqin Wang, Xiaoqi Zhang, Nan Yang, Kai Wu, J. Andrew Zhang, Y. Jay Guo
The paper introduces DHCG, a framework that dynamically constructs hierarchical collaboration graphs for large language model–based multi‑agent systems. DHCG coordinates Planner, Worker, and Generator modules to adaptively determine the composition and scale of agents during execution, guided by feedback and action‑aware preference optimization. Experiments on code generation, mathematical reasoning, and domain‑specific tasks show DHCG surpasses static and dynamic baselines, improving performance by 2.77–8.02 points and achieving a 13.06‑point gain over single‑agent baselines.
By Jie Ren, Jiakang Yuan, Chenyu Huang, Hezeer Ma, Jiayuan Fan, Tao Chen
ShanLiangRen is a nutrition agent designed to generate personalized daily meal plans that satisfy both user constraints and multidimensional nutritional goals. It transforms dietary specifications, nutrient data, user attributes, and natural language requirements into a constrained planning instance, then uses a retrieval‑augmented generation approach to narrow the candidate set from a large ingredient and recipe space. Finally, it refines plans via Pareto‑guided iterative revisions with an LLM, producing fully quantified meal plans with explicit ingredients, portion sizes, and compliance reports.
By Miao Xie, Xiao Zhang, Yuan Wang, Ruixin Zhu, Chunli Lv
The paper introduces a method for sea surface temperature (SST) forecasting that combines textual environmental context with spatial graph representations for large language models (LLMs). Historical SST and anomaly sequences, date‑aligned environmental records, and static ocean knowledge are provided as textual input, while a static graph captures geographic–climatological relations and a dynamic graph captures recent SST correlations and tropical‑cyclone influence. The approach achieves the lowest mean absolute error and highest R² among compared methods over ten forecast steps in the South China Sea, and includes a rule‑based module that links predicted trends to source‑linked contextual explanations.
By Xiong Li, Xiaowei Zhou, Yanwei Yu, Qian Cui, Junyu Dong
The paper introduces Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an LLM agent has successfully completed its task using a single trajectory without requiring privileged model access or training data. CRGs decompose the agent’s overall claim of success into contextualized sub-claims, estimate confidence for each terminal claim based on trajectory evidence, and aggregate these into an overall confidence estimate. Experiments across multiple benchmarks, models, and agent frameworks show that CRGs produce better-calibrated confidence and more effective risk-aware decision making than existing verbalized, sampling-based, and white-box surrogate methods, while also providing transparent, auditable evidence for each estimate.
By Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti, Pouya Pezeshkpour, Estevam Hruschka