arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.
By Umut Onur Yasar
The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.
By Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan
The paper evaluates the common assumption that combining flow statistics and TLS handshake fingerprints improves encrypted command-and-control detection. Using 17,577 TLS flows from 62 real Cobalt Strike captures, the authors show that data leakage and preprocessing choices inflate performance metrics, revealing that the true benefit of multi-view fusion is minimal (0.022 F1). They also uncover that many captures contain only benign traffic and that class imbalance is an artifact of analysis rather than a real feature of the task.
By Hoang-Huy Nguyen-Huu, Van-Tri Phan, Khuong Nguyen-An
arXiv:2606. 12171v1 Announce Type: cross Abstract: Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs.
By Jos\'e Medina, Paul Honeine, Abdelaziz Bensrhair, Amnir Hadachi
arXiv:2607. 15467v1 Announce Type: new Abstract: Knowledge distillation enables an adversary to replicate a proprietary classifier by querying its prediction interface and training a surrogate on the returned probability vectors.
By Khawaja Abaid Ullah, Mohammad Javad Khojasteh
arXiv:2609.36426v1 Announce Type: cross
Abstract: A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhil...
By Trung Minh Bui, Jongsul Moon, YoungOuk Kim, Jung-Hoon Hwang, Dongin Shin
The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.
By Chencheng Zhu
arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.
By Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min
TwinMark is a watermarking scheme that embeds a single SHAKE128 secret into a vision model using two complementary linear functionals of model-output summaries: a covariance projector (cov‑Feat) and a class‑conditional Fisher‑aligned linear carrier (cc‑FALC). These readouts cover both classifier APIs attacked by KL knowledge distillation and representation‑only hosts attacked by feature‑matching distillation, each providing a teacher‑measurable a posteriori certificate. Across 13 attacks on datasets such as CIFAR‑10, CIFAR‑100, and Mini‑ImageNet, TwinMark remains detectable on every post‑attack model that retains task utility, survives cross‑architecture distillation onto ResNet‑18/50, VGG‑16, and MobileNet‑V3, and can be ported to GNSS few‑shot, VOC detection, ISIC segmentation, and STL‑10 SimCLR.
By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv:2609.01345v1 Announce Type: new
Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natura...
By Dushyant Rajput
arXiv:2607. 07050v3 Announce Type: replace-cross Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support.
By Jiabin Shen, Guang Chen, Chengjun Mao
The study examines whether accuracy gaps between source and target tasks can certify the failure of scalar recalibration maps for large language model judges. Across thirteen judges, two generators, eight domains, and 1,176 transfers, the accuracy gap only provides a weak lower bound on target calibration error and can predict opposite outcomes. Even with a finite‑sample lower certificate, the method shows low power (0.13) and does not reliably indicate when recalibration will fail.
By Fariya Afrin, Ibne Farabi Shihab