arXiv AI

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

arXiv Machine Learning
Jul 7

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

arXiv:2602. 20629v3 Announce Type: replace Abstract: As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation.

By Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye, Edward Zhang, Ruslans Aleksejevs, Todor Anti\'c, Polina Baron, Sujeet Bhalerao, Shubhrajit Bhattacharya, Zachary Burton, John Byrne, Hyungjun Choi, Nujhat Ahmed Disha, Koppany Istv\'an Encz, Yuchen Fang, Robert Joseph George, Ebrahim Ghorbani, Alan Goldfarb, Jing Guo, Meghal Gupta, Stefano Huber, Annika Kanckos, Minjung Kang, Hyun Jong Kim, Dino Lorenzini, Levi Lorenzo, Tianyi Mao, Giovanni Marzenta, Ariane M. Masuda, Lukas Mauth, Ana Mickovic, Andres Miniguano-Trujillo, Antoine Moulin, Wenqi Ni, Tomos Parry, Kevin Ren, Hossein Roodbarani, Mathieu Rundstr\"om, Manjil Saikia, Detchat Samart, Rebecca Steiner, Connor Stewart, Dhara Thakkar, Jeffrey Tse, Vasiliki Velona, Yunhai Xiang, Sibel Yal\c{c}{\i}n, Jun Yan, Ji Zeng, Arman Cohan, Quanquan C. Liu
arXiv AI
1d ago

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

The paper reports a specialized training pipeline for large language models to excel in competitive programming, combining problem curation, synthetic reasoning traces, supervised fine‑tuning, and reinforcement learning. Using 22,000 curated problems, the authors train two models—Nemotron‑3‑Nano‑CC (30B) and Nemotron‑3‑Ultra‑CC (550B)—and introduce GenCorrect, a test‑time refinement strategy. On the IOI 2025 benchmark, Nano‑CC scores 468 points with GenCorrect, surpassing the gold‑medal threshold, while Ultra‑CC reaches 502; in IOI 2026, a competition‑specific Ultra‑CC system scores 535.4, exceeding both the gold threshold and the top human score of 498.27, marking the first AI system to outscore the highest‑scoring human contestant on an IOI problem set.

By Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, Boris Ginsburg
arXiv Machine Learning
Jul 30

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

arXiv:2604. 18584v2 Announce Type: replace-cross Abstract: Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity.

By Shaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton, Navid Safaei, Sultan Albarakati, William T. Freeman, Antonio Torralba