← Back to all news
arXiv Computer Vision October 1, 2026 By Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan

Reliability-Aware Checkpoint Selection for Domain Generalization

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
4d ago

From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

arXiv:2609.36569v1 Announce Type: cross Abstract: Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fix...

By Yupeng Chang, Wenxuan Zhang, Yuan Wu
fine-tuning
More like this →
arXiv AI
Aug 14

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

arXiv:2605. 18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy.

By Qinwu Xu, Zhuoheng Li, Jessie Salas
llmsagentscomputer-visionmultimodal
More like this →
arXiv AI
Jul 23

Test Case Prioritization for DNNs via Neural Collapse Instability

arXiv:2607. 20046v1 Announce Type: cross Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important.

By Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin
safety
More like this →
arXiv AI
Jul 24

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

arXiv:2607. 21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios.

By Jingyuan Li, Xiaoyi Jiang, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
diffusion
More like this →
arXiv AI
4d ago

Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

arXiv:2609.37169v1 Announce Type: cross Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since add...

By Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou
llmssafety
More like this →
arXiv Machine Learning
Sep 15

Certifying Model Upgrades with Slice-Wise Non-Regression and Incumbent Fallback

arXiv:2609.13714v1 Announce Type: new Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...

By Shengwei Zhang, Tao Wu, Fei Qian
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea