CONSISTRE is a consistency‑aware framework for document‑level relation extraction that tackles contradictions in large language model predictions. It offers two tracks: an inference‑time track that refines black‑box LLM outputs through constraint‑aware prompting, verification, and self‑reflection, and a training‑time track that distills consistency knowledge into smaller open‑source models via supervised fine‑tuning and reinforcement learning. Experiments on DocRED show both tracks outperform baselines, with the inference‑time track matching competitive F1 scores and the training‑time track narrowing the performance gap to proprietary LLMs while reducing inference cost.
By Mingxuan Sun
Rufus-Air is an open, reproducible post‑training recipe for the GLM‑4.5‑Air‑Base model, structured as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction‑Following RL, General Agent, Coding Agent, Search Agent, and RLHF. The paper documents the data, reward design, infrastructure, stage order, and stagewise results required to reproduce the recipe, emphasizing that stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge‑based signals. Key findings highlight the importance of diverse, high‑quality SFT, difficulty filtering, reward reliability for stage ordering, and the role of infrastructure and engineering choices.
By Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Zixuan Zhang, Tuo Zhao
The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, defining it via a continuation‑conditioned value of a constrained learning‑state intervention and observation‑relative fibers. It presents an executable finite‑frame classifier‑head that protects current logits while repairing historical margins, and distinguishes local admissibility, intervention value, and complete‑policy performance. Experiments on CLINC‑derived class‑incremental tasks, output distillation with RoBERTa, and SGDW dynamics demonstrate that coupling can produce both positive and negative interactions, and that coordinated content controls can match or exceed development gains while guided allocation reduces cross‑entropy loss compared to standard replay.
By Qinyou Wang
ComplexSync is a diffusion-based framework that delivers real‑time, high‑fidelity lip synchronization even in complex scenarios. It uses a dual‑stream joint training strategy to prevent reference‑frame leakage, a distillation‑based acceleration for single‑step denoising that reaches over 70 FPS, and a relational alignment loss that incorporates structural priors from Vision Foundation Models to improve robustness. The authors also introduce the first benchmark for complex lip synchronization, featuring more than 200 challenging video sequences and specialized metrics, and show that ComplexSync outperforms existing methods on both standard and complex tasks.
By Jiaran Cai, Xingpei Ma, Shenneng Huang
The paper introduces a byte‑constrained cooperative perception framework that balances dense coverage with sparse refinement. Each vehicle sends a highly compressed coarse Bird’s‑Eye‑View (BEV) layer covering the entire map and uses the remaining bandwidth to transmit high‑resolution patches selected by a Task‑Aware Benefit Selector. Experiments on DAIR‑V2X and OPV2V demonstrate that this coverage‑refinement strategy achieves superior accuracy‑payload trade‑offs, reaching 0.60 AP@0.7 with only 1.87 KB per non‑ego agent.
By Melih Yazgan, Timon M\"uller, J. Marius Z\"ollner
The paper studies the numerical solution of the Beurling‑LASSO (BLASSO) for estimating Gaussian mixture models (GMMs) with unknown numbers of components and unknown diagonal covariance matrices. It introduces a Conic Particle Gradient Descent (CPGD) algorithm that incorporates Riemannian gradient descent to respect the Fisher‑Rao geometry of Gaussian distributions. The authors provide theoretical convergence guarantees, including exponential local convergence under a non‑degeneracy condition related to component separation, and demonstrate through numerical experiments that CPGD is more robust to overspecification of components than the EM algorithm.
By Romane Giard, Yohann De Castro, Roland Denis, Cl\'ement Marteau
The paper introduces Rift, a two‑stage system that reduces the computational load of vision‑language models on satellites by pruning image tiles that do not affect the answer and then applying elastic prefill to limit token usage. By exploiting answer‑invariant token redundancy, Rift cuts energy consumption by 78 % and latency by 69 % compared to exhaustive tiled inference, while boosting accuracy from 45 % to 73 % on LLaVA‑1.5 7B running on a Jetson AGX Orin.
By Ishani Janveja, Davis Zhang, Seoyul Oh, Deepak Vasisht
GHOST-Q evaluates how post‑training quantization affects visual grounding in vision‑language models. The study compares three 8B VLM families across FP16, INT8, and NF4 precisions, pairing predictions to measure how compression redistributes grounding successes and failures. While most quantized variants maintain overall accuracy, several exhibit significant changes in hallucination‑sensitive conditions, and memory savings do not always translate to lower latency.
By Saim Rehman, Muhammad Shafique
PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.
By Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, Hao Wang, Xiaoqing Zheng, Xuefeng Yan, Lei Bai, Ya-Qin Zhang, Shuyi Zhang, Wei-Ying Ma, Dahua Lin, Bowen Zhou, Hao Zhou
The paper introduces a framework for Flow‑Matching Vision‑Language‑Action (VLA) models that allows independent adjustment of backbone depth, action expert depth, and denoising steps. Lightweight Exit Transformers are added at intermediate layers to enable early exits, and a KV Cache synthesis mechanism manages skipped layers so the action expert can exit deeper than the backbone. Experiments on SmolVLA and π0.5 across LIBERO and Meta‑World show that joint tuning of these compute axes reduces latency by 79.2 % and FLOPs by 31.8 %, while improving mean success rate by 5.6 %.
By Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia
The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.
By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
The paper introduces SPARK, a method for privacy‑preserving continual learning that decouples knowledge retention from privacy correction. SPARK freezes the post‑task distribution and then selectively corrects it to reduce the likelihood of sensitive content while maintaining strong performance on current and past tasks. Experiments show that this approach effectively suppresses PII and preserves continual‑learning utility across various settings.
By Shengtao Wen, Yunying Yang, Xiang Chen, Lingbing Guo, Yu Tian, Sheng-Jun Huang
The paper introduces Selective Supervision for Direct-OPD (S$^2$D-OPD), a refinement of Direct On-Policy Distillation that filters out states where the teacher’s policy change is minimal, as measured by the teacher‑reference Jensen‑Shannon divergence. By masking low‑divergence states and keeping only the top 10% of states per response, S$^2$D-OPD improves held‑out accuracy on AIME and HMMT benchmarks across multiple teacher‑student pairs without additional forward passes.
By Yibo Zhao, Zixuan Yang, Yunshi Lan, Xiang Li
TOLA is a diffusion‑based text image super‑resolution method that eliminates iterative image‑text diffusion by using a one‑step latent adaptation framework. It employs a confidence‑weighted text conditioning module to build a reliable semantic condition and a lightweight latent residual correction module to fix structured residual errors, thereby preserving text fidelity. Experiments show TOLA outperforms existing diffusion‑based TSR methods, achieving at least 2.72 dB higher PSNR on the CTR‑TSR‑Test benchmark.
By Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao
The paper introduces the Support Vector Graph (SVG), a graph index for vector search that uses kernel methods to guarantee navigability in both metric and non‑metric vector spaces, such as inner product similarity. It shows that popular indices like HNSW and DiskANN are special cases of SVG, and proposes SVG‑L0, which adds an ℓ₀ sparsity constraint to enforce bounded out‑degree while maintaining computational efficiency.
By Mariano Tepper, Ted Willke
The paper proposes using lightweight, calibrated System One decision models—specifically JEV and Laya—to improve autonomous penetration-testing harnesses that rely on large language models (LLMs). It defines four key decision points (finding adjudication, severity recalibration, agent pruning, and confirmation loops) and presents a NeuroSploit case study showing differences in severity distribution, runtime, and grading when using TypeSafe System One. The authors review existing System One specifications, discuss various RL-based training approaches, and introduce Rave, a domain‑adapted model with a proposed training and evaluation framework.
By Joas Antonio dos Santos Barbosa
The paper introduces Reinforcement Learning with Verifiable Rewards (RLVR) applied to small search agents, specifically training a Qwen3.5-0.8B model with Group Relative Policy Optimization and an interleaved Wikipedia-search tool on the MuSiQue dataset. Experiments varying reward shapes across three seeds show that RLVR can achieve a 3.8‑fold improvement over an untrained baseline, with the best run reaching a 0.352 average exact match. The study finds that the sparse exact‑match reward, standard in larger models, performs poorly for small models, indicating that reward design must be tailored rather than scaled down from large‑model recipes.
By Gaurisankar Jayadas, Aske Plaat, \'Alvaro Serra-G\'omez, Sandheep P
The paper introduces Privileged Self-Practice (PSP), a method that retains privileged information (PI) in the prompt rather than the loss during on‑policy self‑distillation for multi‑turn agents. PSP injects short per‑task instructions from an analyzer model when rollouts fail, sampling again with the instruction in context and training with the unchanged GRPO objective. Experiments on AppWorld and SWE‑bench Verified show PSP consistently outperforms plain GRPO, boosting task‑goal completion by up to 65% and resolved rate by up to 61% across three student models.
By Xingyu Su, Abhishek Kumar, Qing Ping, Youzhi Luo, Jonathan Buck, Zach Zhang, Subramanian Chidambaram, Vinayak Arannil
The paper presents a rapid pipeline for training and deploying machine‑learning models on the WeBe Band, a wrist‑worn wearable device. It automates the creation of hardware‑efficient models, integrates with the Piccolo AI ecosystem, and supports OTA deployment while profiling latency and memory usage. Experimental results show trade‑offs between classical models and lightweight neural networks for real‑time performance on a microcontroller.
By Ehsan Kourkchi, Asmita Asmita, Houman Homayoun, Mahdi Eslamimehr
The paper introduces iCoder-27B, a 27‑billion‑parameter model for RTL design and GPU kernel optimization that is developed through a recursive AI‑led process with minimal human input. Human experts provide high‑level objectives and reusable research skills, while the agent autonomously selects experiments, diagnoses outcomes, and refines training strategies, coordinating SFT, self‑distillation, and reinforcement learning. iCoder outperforms GPT‑5.5 and Claude‑Opus‑4.8 on several benchmarks, demonstrating the feasibility of building frontier‑competitive models with largely automated development.
By Cheng Yang, Jiayang Lyu, Shangyuan Liu, Guibin Zhang, Jiong Lin, Xinlei Yu, Junchi Yan, Shuicheng Yan, Weinan E, Linfeng Zhang, Linfeng Zhang, Qibing Ren