arXiv AI By Yezhou Cheng, Zehua Yang, Bojun Lin

Signed Lexical Confidence for Risk-Calibrated Intent Routing

Read the original on arXiv AI →

The paper introduces a signed lexical gate that combines a sentence classifier’s logit margin with a sparse lexical model’s support for the predicted intent, assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent. This gate retains more information than unsigned lexical confidence or a hard agreement rule and is calibrated via an independent binomial procedure to meet specified risk targets. Experiments on BANKING77, CLINC150, and HWU64 show that the proposed score reduces the area under the risk‑coverage curve by up to 15.8% and increases accepted coverage at low error rates, offering a compact, interpretable confidence enhancement for risk‑calibrated intent routing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

The paper investigates hallucination detection in black‑box large language models by leveraging two accessible signals: semantic entropy, which captures disagreement among sampled response meanings, and token‑level uncertainty derived from log‑probabilities. It introduces a TopK aggregation technique, a hybrid CoCoA method combining uncertainty with semantic dissimilarity, and two supervised approaches—Gated and Stacked—that integrate token and semantic features. Across seven benchmarks and four language models, the supervised Stacked method performs best in many cases, while TopK and CoCoA remain competitive without labeled data, though all methods require careful threshold calibration.

By Urja Pawar, Rajitha Ramanayake, Owen O'Neill, Nabeel Kemal, Abhishek Mandal, Houssem Chatbri, Christopher Martin
arXiv AI
Sep 2

Validity-Aware Jailbreak Evaluation for Large Language Models

The paper introduces SEAV, a verification‑centric framework for evaluating jailbreak attempts against large language models. SEAV decomposes responses into ordered steps and checks both validity and correctness using LLM‑as‑a‑judge and retrieval‑grounded verification. The method reduces false positives by 14.9 percentage points on a strategic‑dishonesty diagnostic and reclassifies 22.1–51.0% of previously successful jailbreaks as invalid across multiple benchmarks.

By Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran