Overcoming the Incentive Collapse Paradox
arXiv:2603. 27049v2 Announce Type: replace-cross Abstract: AI-assisted task delegation is increasingly common, yet human effort in such systems is costly and typically unobserved.
arXiv:2603. 17212v2 Announce Type: replace-cross Abstract: When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation is noisy.
arXiv:2603. 27049v2 Announce Type: replace-cross Abstract: AI-assisted task delegation is increasingly common, yet human effort in such systems is costly and typically unobserved.
arXiv:2607. 20668v1 Announce Type: cross Abstract: TextGrad improves language-model systems by revising text from feedback.
arXiv:2604. 17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions.
arXiv:2606. 04261v1 Announce Type: new Abstract: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback.
arXiv:2606. 04507v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability.
arXiv:2607. 20083v1 Announce Type: cross Abstract: Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models.
AutoVerifier is a residual‑guided, non‑parametric optimization framework designed to improve reference‑based answer verification. It learns verifier inductive biases from recurring errors, records them as rule cards, and promotes them to code modules or prompt guidance only after replay validation ensures no regressions. Experiments on four verifier benchmarks show that AutoVerifier surpasses state‑of‑the‑art verifiers by a large margin.
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning.
arXiv:2605. 04356v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards has been used to elicit impressive performance from language models in many domains.
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skil...
CHAI for LLMs is a framework that improves large language models’ performance on code‑mixed translation tasks by using LLMs as annotators to create preference data, applying reinforcement learning from AI feedback, incorporating LLM‑generated domain knowledge for iterative refinement, and evaluating on real‑world datasets. The approach yields a 68.45% average win rate over state‑of‑the‑art open‑source models in human‑adjudicated tests. It demonstrates a scalable method to enhance code‑mixed language understanding in open‑source LLMs.