A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.
By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil
arXiv:2607. 08268v1 Announce Type: new Abstract: High-volume structured extraction pays a large model's latency on every item, so distilling the task into a small on-device model is attractive: comparable output at a fraction of the time and cost.
By Vinay Kumar Chaganti
The paper investigates correctness‑gated multi‑teacher distillation, comparing a weighted arm to unfiltered distillation across eight experimental arms. While the weighted arm shows modest gains in accuracy (+0.1660) and macro‑F1 (+0.1323) and a reduction in unsafe action rate (−0.4979), it also exhibits lost label functionality, such as zero Refuted recall and over‑assignment of NotEnoughInfo. A subsequent grounding audit was inconclusive, failing to demonstrate a clear improvement in evidence grounding or overall system performance.
By Xiaofei Feng
arXiv:2609.37616v1 Announce Type: new
Abstract: Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on th...
By Abhinav Rajeev Kumar (Lossfunk), Paras Chopra (Lossfunk)
arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?
By Ely Hahami, Ishaan Sinha, Lavik Jain