arXiv AI

Constraint-Guided Enterprise Data Mapping with Large Language Models

The paper introduces Constraint‑Guided Enterprise Data Mapping (CGM), a neuro‑symbolic approach that uses schema‑grounded admissibility constraints to steer large language models (LLMs) in aligning enterprise data. CGM operates in three stages: defining constraints with metadata, generating candidates under relaxed constraints to ensure feasibility, and ranking them with a bounded LLM. Experiments show that hard constraints dramatically reduce candidate space and improve F1 scores, enabling small models to match or surpass large LLMs at a fraction of the cost while reducing expert effort.

arXiv AI
Aug 3

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

arXiv:2607. 29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree.

By Penglin Zhu, Jungang Xu
arXiv AI
4d ago

CRASM-Gate: Deterministic-First Constraint- and Role-Aware Semantic Mapping with Selective Model Assistance Across Heterogeneous Industrial Standards

The paper introduces CRASM, a deterministic, constraint‑ and role‑aware semantic mapping framework for aligning engineering concepts across incompatible industrial standards, and its extension CRASM‑Gate, which optionally employs a large language model through a confidence gate while preserving deterministic validation. The framework decomposes the mapping process into standard‑specific canonicalization, bounded retrieval, deterministic rules, role interpretation, ranking, ambiguity refusal, and target validation. Experiments on 14,400 sample decisions across six standard pairs show CRASM‑Gate achieving a mean F1 of 0.9938 and perfect structural validity, outperforming a model‑only baseline and reducing latency significantly compared to generative‑model‑only approaches.

By Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
arXiv AI
Aug 26

Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA

The paper introduces Constrained Entity Selection under Partial Knowledge (CES-PK), a framework for improving large language model (LLM) based knowledge graph question answering (KGQA) by filtering candidate answers with lightweight symbolic constraints instead of full semantic parsing. CES-PK uses a three-valued constraint semantics—satisfied, violated, unknown—to handle incomplete knowledge graphs and avoid incorrect rejections under open‑world assumptions. Experiments on the Hetionet biomedical knowledge graph show that applying type, relation, and exclusion constraints increases precision while preserving recall, and that satisfied constraints can be used to rank remaining candidates.

By Emanuel Kitzelmann
arXiv AI
Sep 12

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is a benchmark that evaluates how well large language models (LLMs) understand and apply version-constraint resolution semantics, such as determining whether a version satisfies constraints like ^1.2.3 or >=2.0. The study finds that many models struggle with certain corner cases, with GPT‑5.1 performing poorly while Claude and Opus perform much better. The authors suggest that the failures stem from an activation/application gap rather than a lack of knowledge, and recommend that coding agents delegate version resolution to a dedicated resolver tool.

By Qibai Chen, Zeming Liu
arXiv AI
Aug 5

IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation

arXiv:2608. 02641v1 Announce Type: cross Abstract: Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incorrect objectives, while iterative repair, search, and multi-agent workflows increase inference cost.

By Penglin Zhu, Linhai Zhang, Jungang Xu, Xinchi Wei, Xiuqi Wu
arXiv Computation and Language
Sep 25

CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models

CONSISTRE is a consistency‑aware framework for document‑level relation extraction that tackles contradictions in large language model predictions. It offers two tracks: an inference‑time track that refines black‑box LLM outputs through constraint‑aware prompting, verification, and self‑reflection, and a training‑time track that distills consistency knowledge into smaller open‑source models via supervised fine‑tuning and reinforcement learning. Experiments on DocRED show both tracks outperform baselines, with the inference‑time track matching competitive F1 scores and the training‑time track narrowing the performance gap to proprietary LLMs while reducing inference cost.

By Mingxuan Sun
arXiv Computation and Language
Sep 11

Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization

The paper introduces a Database Normalization Benchmark (DNBENCH) with 3,275 samples to evaluate how well Large Language Models (LLMs) can perform database normalization from 1NF to BCNF, assessing semantic equivalence, structural accuracy, and logical validity. It identifies common failures in dependency inference, schema decomposition, and inter-table constraint reconstruction across various complexity levels. The authors also propose a Multi-Agent Reasoning for Schemas (MARS) framework that separates evidence extraction, violation diagnosis, and decomposition planning from schema generation, achieving an 82.0% improvement in DNB-SCORE over a single-prompt baseline.

By Dong-Jae Koh, Huisu Kim, SeongHwan Yoon, Lasse M. Jantsch, Chun-Hee Lee, Seonghyeon Lee, Young-Kyoon Suh
arXiv Computation and Language
Sep 17

English Word Sense Disambiguation in 2026: When the Labels Become the Bottleneck

The paper reports that in English all‑words word sense disambiguation (WSD), the scarcity of high‑quality labels—not the models—has become the limiting factor. The authors introduce lexEN, a human‑adjudicated correction layer over the Maru2022 ALL_NEW benchmark, and SenseBench, a living leaderboard for LLM WSD evaluation. They show that frontier large language models reach about 95 % accuracy on lexEN‑v1, that relabeling corpora with these models improves downstream systems, and that fine‑grained WordNet senses are often ill‑posed, with coarsening improving both annotator agreement and model performance. "whyItMatters":"The study highlights that improving label quality and managing annotation costs are now the critical challenges for advancing WSD performance, as model accuracy is already near its theoretical ceiling."

By Vassili Philippov, Amro Salman, Dmitrii Andreev, Penny Hands, Emil Kaiumov, Pavel Katunin, Anton Nikolaev