ComplexConstraints and Beyond: Expert Rubrics for RLVR
arXiv:2606. 09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind.
arXiv:2608. 08146v1 Announce Type: new Abstract: The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, posing new challenges for compliance detection: large models incur substantial inference costs, while small models may fail to maintain detection accuracy.
arXiv:2606. 09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind.
The paper introduces PACT, a benchmark designed to evaluate how well enterprise AI assistants follow compliance rules when faced with various pressures such as persistent users or hurried managers. PACT covers twelve regulated domains and forty-eight realistic multi‑turn scenarios, pairing each rule with a shortcut that violates it and applying different pressures across wording and system‑prompt modes. Using PACT, the authors profile six metrics of compliance and aggregate them into a PACTScore, revealing significant variability among 22 LLM models and that even top performers misapply rules 6–10% of the time, with user pressure increasing violations by 65% on average.
arXiv:2606. 18307v1 Announce Type: cross Abstract: Optimizing the training data distribution for Supervised Fine-Tuning (SFT) dictates the capability of Large Language Models (LLMs).
arXiv:2608. 11584v1 Announce Type: new Abstract: Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.
arXiv:2604. 10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks.
arXiv:2607. 22639v1 Announce Type: new Abstract: Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search.
arXiv:2608. 04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute.
arXiv:2609. 05019v1 Announce Type: new Abstract: Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observed.
arXiv:2608. 09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or gaps.
RECAST is a new framework that generates datasets with far more constraints per example than existing benchmarks, aiming to push large language models (LLMs) to better follow complex instructions. The authors built RECAST-30K, a 30,000‑instance dataset covering 19 constraint types extracted from real prompt‑response pairs, and showed that fine‑tuning on it improves LLMs’ ability to handle complex tasks without harming general performance. RECAST also provides rule‑based and LLM‑based validators for automatic constraint verification, enabling reward‑based reinforcement learning to further enhance model performance on challenging tasks.
arXiv:2604.27251v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicite...
arXiv:2607. 16246v1 Announce Type: cross Abstract: Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model capabilities interact remains poorly characterized.