arXiv Computation and Language By Zhaolu Kang, Yingjie He, Kehan Jiang, Leqi Zheng, Jiachen Qian, Qianyuan Zhang, Chunlei Meng, Yujie Feng, Yuan Wang, Stephen Dou, Aming Wu, Pengxiang Zhao, Jiaxin Liu, Guansu Wang, Zeyu Zhang, Lei Wang, Qishi Zhan, Xiaomin He, Meisheng Zhang, Jianyuan Ni, Richeng Xuan

How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

arXiv AI
Aug 18

A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs

arXiv:2510. 06039v2 Announce Type: replace-cross Abstract: Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language text with verifiable Knowledge Graph (KG) facts.

By Chengwei Wu, Xingrui Zhuo, Mingyang Gao, Xinghe Cheng, Zhichao Yan, Jiapu Wang
arXiv AI
Sep 11

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

The paper presents an end‑to‑end pipeline for translating natural language planning descriptions into PDDL problem instances using large language models. It incorporates multiple checks—syntactic parsing, planner success, domain conformance, an LLM critic, and iterative repair—to ensure faithfulness to the original task. Experiments on Planetarium, AutoPlanBench, and curated PDDL~2.1 problems reveal that operational success can diverge from benchmark‑reference reconstruction, and that structured repair improves outcomes while PDDL~2.1 remains challenging for reference reconstruction.

By Joana Rosa, Pedro Santos, Valdemar Oliveira, Rom\~ao Silva, L. Miguel Silveira, Bruno Martins
arXiv Computation and Language
Sep 4

Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

The paper introduces IndicReStruct, a benchmark dataset comprising two variants—GSM8K-Reordered and GSM8K-Voice—derived from GSM8K to test multilingual LLMs on Hindi and Malayalam. It evaluates six state‑of‑the‑art LLMs under constrained constituent reordering and active‑passive voice transformation, finding consistent and significant drops in mathematical reasoning performance. Qualitative error analysis and residual‑stream activation patching reveal that failures often stem from disrupted entity‑quantity alignment, with intermediate transformer layers playing a key role in restoring reasoning.

By Karthika Nhayakkat, Rajat Verma, Maharaj Brahma, Vetcha Gnana Mahesh, Maunendra Sankar Desarkar, Ganesh Ramakrishnan, Rohit Saluja
arXiv Computation and Language
Sep 16

EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting

EviSI is an evidence‑based evaluation agent for low‑latency simultaneous speech‑to‑speech translation. It combines Multidimensional Quality Metrics with interpreter‑developed criteria, using shared source evidence to assess four dimensions—Anchor, Event, Logic, and Fluency—while deduplicating verified errors before scoring. On English‑to‑Chinese and Chinese‑to‑English data, EviSI’s rankings correlate strongly with human judgments, outperforming BLEU and COMET, and its multilingual extension maintains these correlations across five language directions.

By Ben Yan, Zongyao Li, Xiaoyu Chen, Daimeng Wei, Weidong Liu, Huan Zhao, Chong Li, Yaode Wang, Yuzhe Shang