arXiv AI By Minji Kim, Jihyoung Jang, Hyounghun Kim

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Read the original on arXiv AI →

The paper introduces KoNA, a benchmark designed to evaluate selective non‑compliance in vision‑language models (VLMs) across five categories—False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety. KoNA tests both query‑level and component‑level non‑compliance using paired single and compound queries, revealing that many VLMs struggle to refuse, correct, or abstain appropriately, especially when selective non‑compliance is required. Fine‑tuning VLMs on KoNA examples improves non‑compliance accuracy while preserving performance on fully answerable tasks, indicating that models can better distinguish answerable components from those needing non‑compliance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.

By Minji Kim, Hyounghun Kim
arXiv AI
Jun 26

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

arXiv:2606. 26101v1 Announce Type: cross Abstract: Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior.

By Renwei Meng, Bowen Zhang, Jian Wang, Xican Wang, Haoyi Wu, Xuanyan Qiu, Shengan Yang