The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs
Read the original on arXiv Machine Learning →arXiv:2607. 02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpoints of one release line as if the model behind it had not shifted.
Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.