Hugging Face Trending Papers

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

Read the original on Hugging Face Trending Papers →

Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the compressed and original network's internal representations under random probe inputs. This stack has a blind spot.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
Sep 18

Faithful, Not Corrective: Model Capability Governs Message-Format Effects in Multi-Hop Agent Relays

The study investigates how different message formats affect the fidelity of information as it passes through multiple LLM agent relays. Using a controlled testbed, the authors encode twelve atomic facts in five formats (free natural language, precision‑instructed NL, JSON, triples, key‑value) across six hops and evaluate recall against programmatic ground truth. Results show that strong relays maintain near‑lossless recall for all formats, while weaker relays exhibit significant format‑dependent recall loss, and that any injected error is faithfully propagated across all formats without causing collateral damage.

By Sicheng Zeng