arXiv Machine Learning

The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

arXiv:2607. 02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpoints of one release line as if the model behind it had not shifted.

arXiv AI
Sep 3

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

The paper investigates how persistent memory in AI agents can lead to over‑trust in stale facts, creating a "Memory Trust Gap" that worsens as model capability increases. Using a benchmark with Benefit and Safety suites across Qwen3 models of varying sizes, the authors show that larger models are more prone to harmful over‑trust, especially when metadata is absent or misleading. They also demonstrate that mitigation strategies such as exposing metadata or pre‑resolving conflicts improve accuracy, but the effectiveness depends on model size and dataset.

By Jundong Hu, Shekar Ramachandran
arXiv Computation and Language
Sep 1

Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study of Black-Box LLM API Fingerprinting

The study evaluates whether prompt‑token counts can reliably identify the lineage of large language models served via APIs. Using a frozen‑threshold approach on 24 labeled endpoint pairs, the authors find that token‑count consistency perfectly separates development pairs but only half of the holdout pairs meet the strict repeatability criteria, yielding moderate accuracy and perfect specificity. The results confirm token‑count consistency as a fingerprint of shared tokenization stacks but reject it as a standalone test for model‑family attribution.

By Bo Chen
arXiv AI
Sep 17

Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.

By Vishnu Bindu Balachandran
arXiv AI
Sep 17

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

AutoTuneBench introduces a trustworthy measurement protocol for evaluating how large language model agents auto‑tune GPU kernels and serving engines. The benchmark addresses four failure modes—strawman baselines, machine‑dependent timing, saturated tasks, and infrastructure defects—by enforcing code‑frozen protocols, database validation, anti‑cheat checks, pre‑registered comparisons, and external result anchoring. Using this protocol, the authors demonstrate that previously reported speedups are inflated, revealing more modest improvements across different engines and machines.

By Li Chen
arXiv AI
Aug 28

Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems

The paper investigates how approvals issued by a large language model (LLM) guardrail for self‑adaptive systems can become stale between the time of check and the time of use, creating a TOCTOU hazard. It introduces three metrics for verdict freshness, evaluates them across five SAS environments, and proposes the Freshness‑Bounded Shield (FBS) to estimate an approval’s validity horizon without a plant‑dynamics model, reducing expiry rates significantly. The study also audits LLM judges and formulates a freshness contract requiring approvals to remain valid at use time.

By Ilai Shraga, Roei Eshel, Lior Gorelik
arXiv AI
Aug 19

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

The paper introduces ontological trust, a task‑conditioned property of trajectory prefixes, and presents RGE, an online monitor that decomposes trust into Role, Goal, and Evidence. RGE uses LLMs only for structured task and step representations, while trust updates and interventions are deterministic, producing a replayable and auditable trust trajectory. Evaluated on a cross‑domain corpus, RGE outperforms rule‑, judge‑, and shield‑style baselines, achieving over 93% Drift F1 and maintaining high benign coverage.

By An He, Yao Wang, Haibin Zhang
arXiv Machine Learning
Aug 20

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

ProxyGuard is a new method for assessing the reliability of randomized data release mechanisms that use shared target sets. It offers two modes: named-release mode, which corrects for multiplicity and certifies specific releases, and direct shared-target mode, which evaluates independent mechanism draws on a common target, providing finite-sample reliability guarantees without needing independent target batches. In a registered study, direct mode increased power from 5.6% to 64.2% at a 0.95 reliability level, while named mode performed better under high-signal evidence.

By Dipesh Tharu Mahato, Pramod Dhungana
arXiv AI
Sep 16

LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

The paper introduces LSREP, a Longitudinal State‑Replay Evaluation Protocol designed to assess how conversational memory evolves over time, incorporating ordered replay, lifecycle schedules, repeated probes, evolving reference answers, and mechanism‑fidelity checks. It applies LSREP to ICE v2, a local‑first memory middleware, and reports that on three ordinary‑density datasets ICE v2 achieves near‑zero mean quality difference from vector‑RAG while using fewer fragments but slightly more prompt tokens, yet fails catastrophically on a dense dataset. In a public diagnostic, ICE v2 underperforms pure vector‑RAG on LongMemEval, revealing significant multi‑session and temporal failures and a quality‑cost trade‑off rather than superior efficiency.

By Deepesh Sonar