How Inference Compute Shapes Frontier LLM Evaluation
arXiv:2606. 17930v1 Announce Type: new Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving.
The paper introduces Thomson, a frontier AI model developed through continual learning on open-weight models, aiming to democratize access to high-performance AI. It argues that institutions with limited resources can achieve frontier-level performance by applying a modern mid- & post-training stack, preserving model plasticity and stability while minimizing high-impact interventions. Thomson demonstrates competitive performance across agentic tasks, safety, legal, tax, multilingualism, and large-scale deep research, exhibiting a distinctive π-shaped improvement pattern and effectively mitigating the forgetting problem seen in narrow domain adaptation.
arXiv:2606. 17930v1 Announce Type: new Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving.
arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.
arXiv:2606. 00047v1 Announce Type: cross Abstract: Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function of the compute and data used during training.
arXiv:2607. 21482v1 Announce Type: new Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models.
arXiv:2607. 00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget?
arXiv:2608. 13940v1 Announce Type: new Abstract: AI research agents (AIRA) can now propose, implement, and evaluate their own machine learning experiments, but progress on frontier tasks is throttled by cost: a candidate solution can be written in minutes, whereas evaluating it can take hours to days of GPU time.
arXiv:2608. 06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibration, and Multimodal Instruction Tuning.
arXiv:2606. 07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.
FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.
The paper investigates the limitations of post-training AI agents that can autonomously train large language models. It distinguishes between execution-level capability—making adjustments within a chosen training strategy—and strategy-level capability—revising the overall approach based on new evidence. Analysis of many public post-training runs shows that agents lock into a strategy early and then only perform local tweaks, regardless of task. Experiments with experience scaffolds, human guidance, and extra compute improve execution but do not enable strategy reevaluation, indicating that agents lack a mechanism to spontaneously reassess their strategy during training.
Feyospace‑v1 describes a data‑centric framework that enables a seven‑person team to train open‑weight cyber agents, addressing bottlenecks such as environment cost, supervision, and teacher access. The system combines five complementary components—Choulea, SkyReal, Hongzwang, PSBreakup, and Kreator—to analyze reasoning, reduce sampling costs, bypass API limits, restore merged model capabilities, and convert expert interventions into trainable signals. Using a diverse data engine, the team produced 164,269 verified trajectories for supervised fine‑tuning, achieving an average 23.76% improvement on CyberGym and 10.49% on pooled CTF suites, with the final checkpoint ranking 10th on the CyberGym leaderboard and first among comparable‑scale models.
arXiv:2606. 25198v2 Announce Type: replace Abstract: Autonomous AI Research promises to accelerate the scientific progress of machine learning.