arXiv Machine Learning By Junghyun Min, Huseyin Uzunalioglu, Mohamed Trabelsi

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Read the original on arXiv Machine Learning →

The paper investigates how fully autonomous machine‑learning research systems can tackle open‑ended, industry‑grade problems, using telecom ticket retrieval as a case study. It finds that while autonomous research excels at hyperparameter tuning, it lacks human intuition and creativity, yet can achieve about 90% of state‑of‑the‑art performance in a fraction of the time and at modest cost. The authors recommend a hybrid approach where human researchers collaborate with autonomous frameworks for optimal results.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 17

AI Research Preference Models

arXiv:2608. 13940v1 Announce Type: new Abstract: AI research agents (AIRA) can now propose, implement, and evaluate their own machine learning experiments, but progress on frontier tasks is throttled by cost: a candidate solution can be written in minutes, whereas evaluating it can take hours to days of GPU time.

By Thomas Simon Foster, Bassel Al Omari, Tingchen Fu, Thomas Mann, Carl Domond, Lucia Cipolina-Kun, Bhavul Gauri, Muna Aghamelu, Alexander D. Goldie, Eryk Helenowski, Jean-Christophe Gagnon-Audet, Alberto Pepe, Saba Nazir, Daniel Izcovich, Noam Levi, Rishi Hazra, Karen Hambardzumyan, Nicolas Baldwin, Xian Li, Martin Josifoski, Paris Giampouras, Masoud Jalili Sabet, Anya Sims, Hela Momand, Tatiana Shavrina, Despoina Magka, Jason Weston, Yulin Wang, Anirudh Goyal, Jo\~ao Henriques, Yoram Bachrach, Emily McMilin, Jakob Nicolaus Foerster
arXiv AI
2d ago

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

ScholarCatalyst is a new benchmark that evaluates how well AI systems can retrieve research papers that inspire new work. The dataset was created by having 184 lead authors of 207 recent computer science papers annotate which earlier papers helped their projects, providing detailed rationales. The benchmark tests retrieval from the literature available at the start of a project, revealing that current agentic search and even advanced models like Claude Fable 5.1 perform only modestly better than simple embedding retrieval.

By Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn
arXiv AI
Jun 2

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

arXiv:2603. 19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains.

By An Luo, Jin Du, Xun Xian, Robert Specht, Fangqiao Tian, Ganghua Wang, Xuan Bi, Charles Fleming, Ashish Kundu, Jayanth Srinivasa, Mingyi Hong, Rui Zhang, Tianxi Li, Galin Jones, Jie Ding
arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister