arXiv Machine Learning

J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

J-Miner extracts executable decision knowledge from fine‑tuned language‑model classifiers by aggregating internal signals aligned with named concepts and learning rules over them. The mined rules reproduce up to 98.3% of the original classifier’s decisions and outperform compact rules based on input words by 6.0–29.5 percentage points in behavioral fidelity. Moreover, these rules can be transferred to lightweight student models that use only about 1/24 of the parameters yet retain 99.8% of the source classifier’s accuracy.

arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow
arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
arXiv AI
Sep 10

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

The paper introduces XTF, an explainable token‑level noise filtering framework for fine‑tuning large language models. XTF breaks down token contributions into reasoning importance, knowledge novelty, and task relevance, scores them, and masks gradients of noisy tokens to improve fine‑tuning. Experiments on math, code, and medicine tasks across seven LLMs show up to a 13.7% performance boost over standard fine‑tuning.

By Yuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu, Hongbin Zhou, Lan Tao, Yiming Li, Zhan Qin, Kui Ren
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv Computation and Language
Aug 31

Knowing Before Answering: Decoding Language Models for Reliable RAG

The paper introduces a method for determining whether retrieval-augmented generation (RAG) systems have sufficient, insufficient, or conflicting evidence to answer a question. By training a lightweight linear classifier on hidden activations and attention-derived features from 16 language models, the authors demonstrate that these internal signals reliably predict the adequacy of retrieved documents, outperforming prompting-based baselines and specialized RAG models. Analysis shows that middle-layer hidden states carry the most informative signals for this triage task.

By Syed Mahbubul Huq, Christopher Child, Tillman Weyde, Pranava Madhyastha