← Back to all news
arXiv AI September 30, 2026 By Yupeng Chang, Wenxuan Zhang, Yuan Wu

From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • fine-tuning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 14

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

arXiv:2605. 18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy.

By Qinwu Xu, Zhuoheng Li, Jessie Salas
llmsagentscomputer-visionmultimodal
More like this →
arXiv AI
1d ago

Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

arXiv:2609.37169v1 Announce Type: cross Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since add...

By Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou
llmssafety
More like this →
arXiv AI
Jul 23

Test Case Prioritization for DNNs via Neural Collapse Instability

arXiv:2607. 20046v1 Announce Type: cross Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important.

By Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin
safety
More like this →
arXiv Machine Learning
Sep 10

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

arXiv:2609.06107v1 Announce Type: new Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which do...

By Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang
llmsreinforcement-learningbenchmarks
More like this →
arXiv AI
Sep 10

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

arXiv:2609.08966v1 Announce Type: new Abstract: Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the...

By Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges
fine-tuningbenchmarks
More like this →
arXiv Machine Learning
Aug 12

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

arXiv:2608. 10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment.

By Joyjeet Singh
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea