arXiv:2606. 25039v1 Announce Type: new Abstract: Recovering governing Ordinary Differential Equations (ODEs) from data is a central challenge in modeling dynamical systems across scientific domains.
By Nikhil Abhyankar, Sha Li, Sanchit Kabra, Naren Ramakrishnan, Yulia Gel, Chandan K. Reddy
The paper investigates integer‑sequence benchmarks from the OEIS by applying a two‑part minimum description length (MDL) learner that searches for P‑recursive recurrences. It finds that MDL difficulty correlates with a combinatorial parameter count, that most sequences fit a recurrence on a prefix but not at full length (the “wilderness” regime), and that language models do not hallucinate in the wilderness but instead hedge, showing that memorisation dominates perceived competence. The study provides a cheap, contamination‑free difficulty signal for OEIS‑derived benchmarks.
By Sabilashan Ganeshan
The paper investigates how large language models handle domain-specific jargon, comparing a general-purpose Llama‑3.1 with a version fine‑tuned on medical data. Two new medical jargon benchmarks reveal that the general model actually outperforms the fine‑tuned variant, and interpretability tools show the fine‑tuned model over‑emphasizes a few components linked to jargon predictions. Reweighting these components narrows the performance gap, and some jargon‑sensitive components also aid materials‑science tasks, indicating a partially domain‑agnostic representation of specialized terminology.
By Darin Keng, Zhewei Sun
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao
The paper introduces the problem of cross‑lingual loopholes in large language model (LLM) unlearning, where forgetting a fact in one language can leave it accessible in others. It presents a new 174‑language benchmark, the Cross‑Lingual Unlearning Tensor, and proposes COVER, a method that selects a subset of source languages to maximize unlearning coverage under a language budget. Experiments show COVER reduces residual knowledge by 7.8–27.3% compared to uniform selection and works on both synthetic and real low‑resource news data.
By Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
The paper audits 22 frontier language models on 12 molecular property regression benchmarks to assess verbatim retrieval of published values. It finds widespread but benchmark‑specific retrieval, with over 50% of models retrieving exact values on five datasets and isolated occurrences on others. Experiments at different reasoning levels show that higher reasoning increases retrieval flags, and attempts to interrupt retrieval reveal that top models can still recognize transformed SMILES and original labels. Suppressing retrieval reduces prediction error variance, indicating that predictive performance is not solely due to memorized values.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
arXiv:2607. 04108v1 Announce Type: new Abstract: Large language models are increasingly used as evolutionary engines for scientific discovery: generate candidates, select winners, feed them back as parents, and repeat.
By Pan Li
arXiv:2604. 27540v2 Announce Type: replace Abstract: Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data.
By Chaemin Jang, Woojin Park, Hyeok Yun, Dongman Lee, Jihee Kim
arXiv:2607. 12649v1 Announce Type: new Abstract: Recent work on extractable memorization in LLMs suffers from two contrasting validity problems.
By A. Feder Cooper, Marika Swanberg, Jamie Hayes, Lea Duesterwald, Christopher De Sa, Daniel E. Ho, Mark A. Lemley, Percy Liang
arXiv:2606. 29182v1 Announce Type: new Abstract: Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to test next.
By Dhruv Agarwal, Reece Adamson, Andrew McCallum, Peter Clark, Ashish Sabharwal, Bodhisattwa Prasad Majumder
arXiv:2608. 15669v1 Announce Type: new Abstract: Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs.
By Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang
The paper audits whether synthetic distractors in RLVR corpora act as shortcuts for learning policies. A classifier using only surface statistics barely outperforms chance, and manual inspection reveals that code distractors are almost identical to correct answers. Experiments with a paraphrase‑matched control show no exploitation advantage for the unmodified data, indicating that the detectable artifact was not used by the policy.
By Esther Xin