Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 16416v1 Announce Type: new Abstract: Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can select and tune known components, but cannot produce structure outside that space.
SAGE-Loop is a new closed‑loop, self‑adaptive AutoML framework that uses large language models to generate and validate machine learning pipelines in multiple rounds, allowing trial‑and‑repair and adaptive ensemble selection for both supervised and unsupervised tasks. It addresses the lack of instant feedback and correction in existing AutoML by enabling process‑level recovery from failures and dynamic use of model diversity. Experiments on 20 public datasets show consistent improvements in performance and stability across classification, regression, and clustering, and demonstrate the system’s ability to recover from execution failures.
The paper introduces AutoMIA, a framework that uses large language model agents to automatically design and implement new membership inference attack (MIA) signal computations. By systematically exploring a wide range of attack strategies, AutoMIA discovers novel MIAs tailored to specific target models and datasets, achieving up to a 0.18 absolute improvement in AUC over existing methods. This demonstrates that LLM agents can serve as an effective and scalable approach for creating state‑of‑the‑art MIAs.
arXiv:2607. 12058v1 Announce Type: cross Abstract: Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation.
arXiv:2607. 04729v1 Announce Type: cross Abstract: LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace.
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.