arXiv Machine Learning By Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer, Jeff M. Phillips

Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage

Read the original on arXiv Machine Learning →

The paper examines binary-choice truthfulness benchmarks, showing that systematic differences in surface-level features between correct and incorrect answers allow models to perform well without genuine reasoning. Using a six-feature logistic classifier, the authors demonstrate that such leakage is detectable and exploitable, and they find similar artifacts in multiple benchmarks. To mitigate this, they propose a cleaning method called Audit‑Prune that removes the most leakage‑reinforcing answer pairs, releasing a revised TruthfulQA dataset with reduced surface‑feature leakage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

Efficient Constrained Graph Search for Post-hoc Error Correction in Binary Classifiers

The paper presents a model‑agnostic framework that performs constrained post‑hoc error correction for binary classifiers. It searches for an interpretable conjunction of feature–threshold rules that corrects remaining false positives or false negatives while limiting newly introduced errors, using graph‑based search, depth‑dependent constraints, and a reduced‑histogram threshold evaluation. Experiments on a large binary‑classification problem show that the method can efficiently identify compact correction rules, such as a configuration that removes 90% of false positives while only sacrificing 5% of true positives.

By Qinwu Xu
arXiv AI
Sep 10

SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.

By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv Machine Learning
Sep 21

LLMs as Feature Engineers for Text-and-Tabular Prediction

The paper presents an iterative framework that uses large language models (LLMs) to automatically extract interpretable, schema‑bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic definitions, an extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance, with error‑driven natural‑language feedback guiding the search. Across three public datasets, the error‑driven loop speeds up feature discovery up to three times and the resulting features outperform any subset when combined with TF‑IDF and dense embeddings, while also providing instance‑level interpretability through SHAP importance rankings and a semantic audit trail.

By Merwan Barlier, Blaz Skrlj
arXiv AI
Aug 24

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.

By Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu