arXiv AI By Muhammad Faraz Shoaib, Muhammad Qasim, Raisulhaq Mohammed Rizwan, Rahmatullah Safdar, Muzammil Adnan Shaik, Abdul Aleem Mohammed

LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems

Read the original on arXiv AI →

LOGIC is a benchmark and evaluation framework that tests how language models can ground engineering requests in a deterministic inventory of candidate changes before propagating selected changes through an electrical traceability graph. The benchmark includes 168 scenarios—144 for selection and 24 for abstention—and evaluates three 7–8B models against intent‑agnostic, lexical, and structured‑evidence methods. Results show that structured evidence can achieve perfect candidate F1 on anchored cases, while large language models perform better on relational‑paraphrase cases; however, grounding accuracy drops as candidate inventories grow, and strict evidence gating reduces false positives but may also remove correct selections.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

TRACE is a new benchmark that evaluates the safety of Large Reasoning Models (LRMs) across the entire inference pipeline, including prompts, reasoning traces, and final responses. It provides prompts in two languages covering nine risk categories and ten attack strategies, and for each prompt four LRMs generate traces and responses that are annotated for safety with supporting evidence extracted from the source text. Evaluation of 18 guardrail models on TRACE shows that detecting unsafe content in reasoning traces is much harder than in prompts or final responses, and that current models struggle to extract the necessary evidence.

By Zhenyu Wu, Siyuan Chen, Changchun Yang, Jiaqi Dong, Min Zhou, Ali Almadan, Talal Hammad, Faisal Wahbo, Aminullah Tora, Mona Alshahrani, Xin Gao
arXiv AI
Sep 2

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

SAGE (State‑Grounded, Abstention‑Aware Evaluation) is a new framework for assessing task‑oriented dialogue agents that focuses on whether each turn correctly advances the workflow state rather than just the quality of the reply. It compiles workflow specifications and per‑turn state differences into schema‑grounded criteria, then evaluates them through a cascade of symbolic rules and encoder/NLI verifiers that abstain instead of guessing, producing a turn‑level decision with an evidence trace. In experiments across MultiWOZ, Schema‑Guided Dialogue, and ABCD datasets, SAGE‑Core—using only symbolic rules and on‑device encoders—outperforms all evaluated LLM‑based judges, including GPT‑4.1 variants, while incurring zero paid LLM cost.

By Rayan Khoury, Shih-Yao Lin, Pratyush Mishra