arXiv:2607. 23191v2 Announce Type: replace Abstract: Fine-tuned code LLMs are often conditioned on a design-intent header to steer parametric CAD generation, but whether the model reads that header's content has been tested neither under execution-level scoring nor with a causal control.
By Yang Xiao
arXiv:2607. 23191v3 Announce Type: replace Abstract: Fine-tuned code LLMs are routinely conditioned on a design-intent specification, but the correctness axis of such a signal -- a wrong intent rather than an absent one -- has not been tested, and the benefit of conditioning is usually scored with the same detector that defines the signal.
By Yang Xiao
The paper investigates why reasoning‑augmented text‑to‑image models like GoT‑R1 sometimes fail on compositional prompts. By separating the explicit textual plan from the decoder, the authors show that the decoder faithfully executes the plan while the planner often writes incorrect spatial relations, especially for phrasing‑dependent cues. Editing or replacing the plan improves image quality without retraining, demonstrating the viability of modular planner‑decoder architectures.
By Ashritha Gonuguntla
arXiv:2608. 09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predictable changes in model function.
By Chencheng Zhu, Xiaoyang Li, Taotao Cai
arXiv:2607. 29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree.
By Penglin Zhu, Jungang Xu
The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”
By Liyun Zhang, Jiayi Guo