arXiv Machine Learning By Xian Wu, Haoran Li, Yuanqi Chu, Dongbin Zhao, Bin Wang

Reinforcement Learning with a Bilevel World-Model Architecture for Scan-Order Optimisation in Laser Directed Energy Deposition

Read the original on arXiv Machine Learning →

arXiv:2605. 25063v2 Announce Type: replace Abstract: Scan-order design in laser directed energy deposition (LDED) is a delayed, path-dependent thermo-mechanical decision problem, because sequence quality becomes observable only after the complete deposition and cooling cycle.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
Sep 15

Harnessing human expertise for high-precision robotic assembly in industrialized construction: A sample-efficient installer-in-the-loop interactive reinforcement learning framework

arXiv:2609.13234v1 Announce Type: cross Abstract: Industrialized construction imposes stringent precision requirements on robotic assembly of modular components such as prefabricated window units. In...

By Zekai Jin, Huiguang Wang, Xiaoning Sun, Yi Shao
arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv Machine Learning
Jun 26

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models

arXiv:2510. 09976v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models such as OpenVLA, Octo, and $\pi_0$ have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fundamentally constrained by the quality and coverage of supervised data.

By Mingyang Lyu, Yinqian Sun, Erliang Lin, Huangrui Li, Ruolin Chen, Feifei Zhao, Yi Zeng
arXiv AI
Aug 26

Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling

The paper introduces a reinforcement‑learning‑guided evolutionary policy optimization framework for scheduling heterogeneous agile Earth observation satellites, addressing task selection, satellite assignment, and sequencing under diverse visibility windows, maneuvering constraints, energy use, and storage limits. It combines assignment‑based indirect encoding with decoder‑based cost evaluation to capture satellite‑dependent constraints while integrating task gain, energy savings, and load balance into a single utility metric. The resulting RLOSMEA algorithm uses reinforcement learning to select high‑level search operators, achieving higher weighted utility and more stable convergence than baseline metaheuristics across varied AEOS scenarios.

By He Wang, Junyu Wu, Hui Li, Yanjie Song, Witold Pedrycz, Liang Li