arXiv AI

Reliable Self-Evolution with Imperfect Proxy Rewards

The paper introduces Conformal Interval-Driven Self-Evolution (CISE), a method that uses conditional conformal inference and online density-ratio estimation to create candidate‑specific reward intervals for self‑evolving search in materials science. CISE applies conservative interval‑based rewards, ensuring that only candidates whose required property intervals lie entirely within feasible regions are returned. Experiments on three self‑evolving search tasks show that all candidates returned by CISE are true positives under high‑fidelity evaluation, whereas baseline methods produce more candidates but include false positives.

arXiv AI
Jun 26

The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators

arXiv:2606. 26294v1 Announce Type: cross Abstract: Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains.

By Alex Iacob, Andrej Jovanovi\'c, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccol\`o Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane
arXiv AI
Jun 10

Towards Diverse Scientific Hypothesis Search with Large Language Models

arXiv:2606. 10587v1 Announce Type: cross Abstract: Large language models (LLMs) are on the rise for accelerating scientific discovery, most recently in advanced tasks such as generating valid scientific hypotheses.

By Haorui Wang, Parshin Shojaee, Kazem Meidani, Kunyang Sun, Jos\'e Miguel Hern\'andez-Lobato, Teresa Head-Gordon, Jiajun He, Chandan K. Reddy, Chao Zhang, Yuanqi Du
arXiv AI
Sep 10

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

The paper introduces Feedback‑Enriched Environments (FEEs) as a new approach to training large language models as autonomous agents for long‑horizon tasks. By shifting from action guidance to observation enrichment during later stages of exploration, FEEs improve performance across SciWorld and BFCL benchmarks with various Qwen3 model scales and RL algorithms. The study shows that FEEs stabilize training, promote proactive exploration, embed environmental guidance into policy weights, and highlight intra‑group feedback consistency as key for stable optimization.

By Hongbang Yuan, Zhuoran Jin, Yixin Cao