arXiv Computation and Language By Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

Read the original on arXiv Computation and Language →

OmniConfess is a training‑free method designed to reduce hallucinations in omni‑modal large language models (OmniLLMs) that handle text, images, audio, and video. The approach fixes a candidate response and re‑scores it at token resolution while selectively intervening on evidence from each modality, producing a token‑by‑channel confession that shows which evidence supports each part of the response. Using this confession, OmniConfess preserves grounded content and corrects commitments that rely on irrelevant or contradictory evidence. The authors evaluated the method on OmniHalluBench, a 3,540‑example benchmark drawn from six datasets across multiple modalities and tasks, and found that OmniConfess mitigates hallucinations across diverse settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 15

Omni-Streaming Thinking

arXiv:2609.15128v1 Announce Type: new Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support...

By Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou
arXiv Computation and Language
Sep 11

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.

By Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
arXiv AI
4d ago

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

The paper introduces IMAVB, a 500‑clip benchmark that tests whether omnimodal large language models can detect when a textual premise contradicts their visual or audio input. Experiments on eight open‑source models and Gemini 3.1 Pro reveal a Representation‑Action Gap: internal states encode mismatches, yet the models rarely reject false premises, exhibiting under‑rejection or over‑rejection. A probe‑guided logit adjustment improves rejection behavior, suggesting the main bottleneck is in translating perception to action rather than in perception itself.

By Trung Nguyen Quang, Yiming Gao, Fanyi Pu, Kaichen Zhang, Shuo Sun, Ziwei Liu
arXiv Computer Vision
Sep 3

Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

The paper examines six inference-time hallucination mitigation methods applied to three large vision-language models across four benchmarks, including MMStar. It finds that reducing hallucination rates often comes at the cost of lower informativeness—such as decreased object recall, visual coverage, and response detail—and that gains on hallucination benchmarks do not consistently translate to improved performance on fine-grained perception and reasoning tasks. The authors argue that current evaluation protocols may overstate progress by favoring conservative generation, and propose that hallucination mitigation should be assessed as a trade-off among faithfulness, informativeness, and overall capability.

By Mehrdad Fazli, Sina Mansouri, Mohit Marvania, Ziwei Zhu