The paper examines binary-choice truthfulness benchmarks, showing that systematic differences in surface-level features between correct and incorrect answers allow models to perform well without genuine reasoning. Using a six-feature logistic classifier, the authors demonstrate that such leakage is detectable and exploitable, and they find similar artifacts in multiple benchmarks. To mitigate this, they propose a cleaning method called Audit‑Prune that removes the most leakage‑reinforcing answer pairs, releasing a revised TruthfulQA dataset with reduced surface‑feature leakage.
By Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer, Jeff M. Phillips
arXiv:2607. 06637v1 Announce Type: new Abstract: In this work, we propose a unified approach for diagnosing misclassification and assessing the robustness of black-box classifiers.
By Evgenii Kuriabov, David Miller, Jia Li
The paper presents an exact constrained reformulation for direct metric optimization (DMO) in binary imbalanced classification, focusing on precision, recall, and F1-score under three settings: fixing precision to optimize recall, fixing recall to optimize precision, and optimizing F1-score. Unlike prior approaches that use smooth approximations, the authors introduce exact penalty methods to solve these problems efficiently. Experiments on benchmark datasets show that this exact reformulation and optimization (ERO) framework outperforms state‑of‑the‑art methods for all three DMO tasks.
By Le Peng, Yash Travadi, Chuan He, Ying Cui, Ju Sun
Decision trees generate interpretable if--then rules, yet they contain irrelevant conditions (IRCs). These IRCs arise from the structural mechanism of tree splitting and persist even in modern optimal sparse tree induction algorithms.
The paper introduces a logic-based framework that extracts global logical rules for node classification in Simple Graph Convolution (SGC) networks. It uses minimal abductive explanations—small sets of node-feature pairs that preserve a node’s predicted class—as an intermediate step. Decision trees trained on these explanations yield compact global rules that retain high fidelity to the original SGC model, as demonstrated on benchmark datasets.
By Bryan Lima Cavalcante, Thiago Alves Rocha
arXiv:2508.10148v2 Announce Type: replace-cross
Abstract: Accurate and explainable out-of-distribution (OOD) detection is required to use machine learning systems safely. Previous work has shown that...
By Maria Stoica, Francesco Leofante, Alessio Lomuscio
arXiv:2606. 02326v1 Announce Type: new Abstract: Hard constraints are usually treated as terminal vetoes: once a candidate violates a requirement, the learned rule rejects it and any repair is handled outside the decision semantics.
By Yifan Wang
arXiv:2607. 13874v1 Announce Type: new Abstract: Decision trees generate interpretable if--then rules, yet they contain irrelevant conditions (IRCs).
By Jung-Sik Hong, Jeongeon Lee, Min Kyu Sim, Sangheum Hwang
arXiv:2603. 06952v2 Announce Type: replace Abstract: As graphs scale to billions of nodes and edges, graph Machine Learning workloads are constrained by the cost of multi-hop traversals over exponentially growing neighborhoods.
By Yuhang Song, Naima Abrar Shami, Romaric Duvignau, Vasiliki Kalavri
arXiv:2608.29262v1 Announce Type: cross
Abstract: Decision trees are attractive for tabular prediction tasks because each prediction follows an interpretable sequence of feature-threshold tests. Unde...
By Hanul Park, Jeonghoon Choi, Juseong Kim, Sanghun Sel, Giltae Song
GraphIFE addresses the class imbalance problem in graph-structured data by tackling a quality inconsistency issue in synthesized nodes. The framework uses graph invariant learning to strengthen embedding space representations and identify invariant features, leading to improved performance on minority classes. Experiments show that GraphIFE consistently outperforms various baselines across multiple datasets.
By Fanlong Zeng, Wensheng Gan, Kangjie Chen, Philip S. Yu
arXiv:2606. 31653v1 Announce Type: cross Abstract: Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimising upper bounds on the worst-case loss over an allowed perturbation set.
By Matteo Melis, Jesus Martinez Del Rincon, Vishal Sharma