arXiv AI

Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

arXiv Machine Learning
6d ago

Decoupled and Distilled: Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer for Few-Shot Class-Incremental Learning

The paper introduces TALON, a Task‑Adaptive LoRA‑Teacher framework for Few‑Shot Class‑Incremental Learning. TALON assigns a dedicated LoRA‑Teacher to each incremental task, then distills the frozen teachers into a single LoRA‑Student via Ensemble Knowledge Transfer, using a semantic‑guided weighting scheme to reduce forgetting and overfitting. Experiments on four FSCIL benchmarks show that TALON matches or surpasses state‑of‑the‑art accuracy while using up to 33× fewer deployment parameters and cutting inference time by 41.7%.

By Hongwei Zhao (School of Computer Science,Engineering, Beihang University), Rui Liu (School of Computer Science,Engineering, Beihang University), Yansong Liu (School of Computer Science,Engineering, Beihang University), Zhiyuan Zou (School of Computer Science,Engineering, Beihang University), Yong Chen (School of Computer Science, Beijing University of Posts,Telecommunications)
arXiv AI
6d ago

Distilling Directional Verification

arXiv:2610.00997v1 Announce Type: cross Abstract: Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may...

By Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim
arXiv AI
6d ago

Distilling LLM Reasoning into Graph of Concept Predictors

The paper introduces Graph of Concept Predictors (GCP), a reasoning-aware active distillation framework that captures a large language model’s intermediate reasoning as a directed acyclic graph of concepts and mirrors this structure in a smaller student model. GCP improves sample efficiency by using a graph-aware acquisition strategy that weighs concept uncertainty, gradient diversity, and node centrality, and enhances training stability through targeted sub‑module retraining that updates only the most influential concept predictors. Experiments on eight NLP classification benchmarks show that GCP achieves better performance under limited annotation budgets while providing more interpretable and controllable training dynamics.

By Ziyang Yu, Liang Zhao
Hugging Face Trending Papers
Sep 2

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

SHELF is a Python system that creates controlled benchmark data and evaluation tasks for libraries and archives, using labelled taxonomies, writing specifications, and a generation budget. It generates 62,899 model-written documents based on Library of Congress vocabularies and supports tasks such as classification, clustering, retrieval, pair classification, and instruction retrieval. The release compares various methods—including TF, TF-IDF, BM25, popular encoders, and zero-shot decoders—showing that sparse methods remain competitive on classification and that SHELF can vary bibliographic facets independently while generating new, verifiably unseen documents.

arXiv Machine Learning
Sep 7

Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation

The paper presents a two-level framework for scalable trade‑up recommendation. Level 1 distills large‑language‑model reasoning into a compact, non‑generative student that classifies product pairs using only precomputed embeddings, achieving high AUC on a benchmark. Level 2 applies product‑type test‑time training to fine‑tune lightweight adapters, further improving performance while keeping inference fast and inexpensive.

By Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi
arXiv AI
Sep 4

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

SHELF is a Python system that creates synthetic, controlled benchmark data for evaluating large language models on bibliographic tasks such as classification, clustering, retrieval, pair classification, and instruction retrieval. It generates 62,899 model-written documents based on Library of Congress vocabularies and compares methods like TF, TF‑IDF, BM25, popular encoders, and zero‑shot decoders, reporting performance metrics such as 0.8887 for subject classification and 0.2605 for genre‑form classification. The tool also allows independent variation of bibliographic facets and can produce unseen documents beyond a model’s training cutoff, with results indicating that model rankings transfer more reliably than absolute scores when compared to other benchmarks.

By Michael J. Bommarito II
arXiv AI
Sep 15

LLMs or Naive Bayes? Old Gems or New Ways

The paper compares Complement Naive Bayes (NB) with zero‑shot and few‑shot large language models (LLMs) across a wide range of model sizes and text classification tasks. NB outperforms LLMs when labeled data is available, achieving comparable accuracy to large LLMs while running thousands of samples per second on a CPU. In zero‑data sentiment settings, LLMs still dominate, but NB remains the best choice for resource‑constrained HPC practitioners, and the authors provide a Kubernetes Helm operator to automate model selection.

By Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein
arXiv AI
Jun 24

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

arXiv:2606. 24259v1 Announce Type: cross Abstract: Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge.

By Noor Islam S. Mohammad, Ulug Bayazit