arXiv AI

ZIVARI-TLBO: A Zero-Cost Inter-Group Evaluated-Elite Relay Mechanism for Teaching-Learning-Based Optimization

arXiv:2606. 17087v1 Announce Type: cross Abstract: ZIVARI-TLBO is a grouped Teaching-Learning-Based Optimization (TLBO) method that augments an existing population-state controller with a fixed inter-group evaluated-elite relay.

arXiv AI
Jun 10

Structure from Reasoning, Numbers from Search: On-Premise Open LLMs as Structural Priors for Coupled MIMO Controller Tuning

arXiv:2606. 11015v1 Announce Type: new Abstract: Tuning controllers for strongly coupled multi-input multi-output (MIMO) industrial processes is hard: decentralized classical auto-tuning ignores loop interaction, and local numerical optimization from natural initializations stalls in the resulting non-convex cost landscape.

By Jiaxuan Chen, Haonan Li, Yang Shu
arXiv Machine Learning
Sep 22

Offline Reinforcement Learning for Distribution-Grid Protection

The paper investigates using offline reinforcement learning to improve line‑selective tripping in distribution grids. A convolutional Q‑network trained with conservative Q‑learning (CQL) processes voltage‑current phasor and impedance data, optionally with raw waveforms, to predict faulted lines. On a realistic CIGRE medium‑voltage network, the best model achieved high per‑timestep precision, recall, and F1‑score, and correctly identified the first trip action in over 98% of fault episodes, though it mis‑tripped in a notable fraction of non‑fault cases.

By Julian Oelhaf, Alexander Luce, Christian Bergler, Andreas Maier, Siming Bayer
arXiv Machine Learning
Jul 9

Optimization-Embedded Active Multi-Fidelity Surrogate Learning for Multi-Condition Airfoil Shape Optimization

arXiv:2603. 17057v2 Announce Type: replace-cross Abstract: Active multi-fidelity surrogate modeling is developed for multi-condition airfoil shape optimization to reduce high-fidelity CFD cost while retaining RANS-consistent aerodynamic metrics.

By Isaac Robledo, Alberto Vilari\~no, Arnau Mir\'o, Oriol Lehmkuhl, Carlos Sanmiguel Vila, Rodrigo Castellanos
arXiv Machine Learning
Aug 12

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

arXiv:2608. 10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction.

By Yuhang Yao, Zeyu Wang, Wanyi Chen, Tongyun Yang, Yuhang Han, Jie Xiao, Chengke Bao, Tianyi Zhao, Lynn Ai, Eric Yang, Tianyu Shi
arXiv Machine Learning
Sep 22

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.

By Chencheng Zhu