arXiv AI

SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents

SeisEvo is a method that uses a large language model (LLM) and multi‑agent search to evolve seismic data reconstruction algorithms rather than optimize a single result. Starting from a classical algorithm, the agents modify only user‑opened components, rejecting candidates that violate physical constraints and scoring the rest by execution. The resulting white‑box algorithms—such as a residual‑gated, phase‑aligned dip‑consistency projection for interpolation and a reliability‑grouped singular‑value shrinkage for simultaneous interpolation and denoising—outperform classic methods by several decibels and generalize to unseen data.

arXiv AI
Jul 7

Generative wave propagator

arXiv:2607. 04440v1 Announce Type: cross Abstract: Seismic wavefield simulation is fundamental to seismology, but conventional finite-difference (FD) methods remain limited by numerical dispersion and stability constraints, which often require dense spatial grids and small time steps and thereby severely limit the effectiveness of iterative inversion workflows.

By Shijun Cheng, Tariq Alkhalifah
arXiv Machine Learning
Jun 4

Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs

arXiv:2606. 04360v1 Announce Type: cross Abstract: Symbolic regression (SR) discovers compact mathematical expressions from data, yet recent LLM-based evolutionary methods remain sample-inefficient because they rely mainly on scalar feedback such as MSE.

By Xinyu Pang, Zhanke Zhou, Xuan Li, Fangrui Lv, Shanshan Wei, Sen Cui, Bo Han, Changshui Zhang
arXiv AI
Sep 17

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

The paper introduces DualViewEval, a benchmark compression technique for agent evaluation that jointly uses outcome and process signals to learn an exact-size miniset and predict full-benchmark scores. By analyzing large-scale trajectories, the authors identify six complementary process signals linked to final agent performance. Across five agent benchmarks and five baselines, DualViewEval achieves superior compression (24×–40×) and lower error rates, while also revealing capability differences among agents for efficient development.

By Xinshuai Guo, Junjie Wu, Dolly Deng, Yinghui Li, Hai-Tao Zheng, Suncong Zheng, Maxm Pan
arXiv AI
Jun 4

Can Generalist Agents Automate Data Curation?

arXiv:2606. 04261v1 Announce Type: new Abstract: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback.

By Feiyang Kang, Hanze Li, Adam Nguyen, Mahavir Dabas, Jiaqi W. Ma, Frederic Sala, Dawn Song, Ruoxi Jia
arXiv Machine Learning
Sep 18

Seismic Site Response Prediction from Sparse Observations Using Finite-Element-Pretrained Latent Dynamics

The paper introduces FLARE‑T, a Transfer‑Enabled Forced Latent Autoencoder for Response Equations, which learns low‑dimensional latent dynamics from dense finite‑element simulations and calibrates them with sparse field observations. By mapping simulated sensor responses into a learned coordinate system, FLARE‑T improves multi‑depth acceleration predictions and pseudo‑acceleration spectra, reducing errors across various sensor locations and motion intensities. Evaluation on a layered‑soil centrifuge test and the Lotung field array demonstrates that FLARE‑T achieves comparable accuracy with different source models, indicating less reliance on precise prior calibration.

By Yi Zhu, Su Chen, Xiaojun Li
arXiv AI
Jul 9

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.

By Kabir Moghe, Peter Chin