arXiv AI By Vasiliy Seibert

Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark

Read the original on arXiv AI →

arXiv:2608. 15255v1 Announce Type: new Abstract: Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific domain.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

The paper presents an end‑to‑end pipeline for translating natural language planning descriptions into PDDL problem instances using large language models. It incorporates multiple checks—syntactic parsing, planner success, domain conformance, an LLM critic, and iterative repair—to ensure faithfulness to the original task. Experiments on Planetarium, AutoPlanBench, and curated PDDL~2.1 problems reveal that operational success can diverge from benchmark‑reference reconstruction, and that structured repair improves outcomes while PDDL~2.1 remains challenging for reference reconstruction.

By Joana Rosa, Pedro Santos, Valdemar Oliveira, Rom\~ao Silva, L. Miguel Silveira, Bruno Martins
arXiv Computation and Language
Sep 16

SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP

SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.

By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
arXiv AI
Sep 15

Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

The paper investigates how large language models handle domain-specific jargon, comparing a general-purpose Llama‑3.1 with a version fine‑tuned on medical data. Two new medical jargon benchmarks reveal that the general model actually outperforms the fine‑tuned variant, and interpretability tools show the fine‑tuned model over‑emphasizes a few components linked to jargon predictions. Reweighting these components narrows the performance gap, and some jargon‑sensitive components also aid materials‑science tasks, indicating a partially domain‑agnostic representation of specialized terminology.

By Darin Keng, Zhewei Sun