arXiv AI By Wenhui Chen

Conflicting Supervision Moves Commitment, Not Capability: A 12.29{\sigma} arrangement effect that is exactly zero under a convention-agnostic score

Read the original on arXiv AI →

The paper investigates how the ordering of training data written under two incompatible but correct conventions influences a model’s learned parameters. It shows that the learning‑rate schedule acts as an averaging operator that determines the ordering effect, with constant schedules producing larger effects than decaying ones. Experiments on a single corpus with fixed budget demonstrate a statistically significant 12.29‑sigma shift in parameter commitment, yet this shift is zero when measured with a convention‑agnostic metric.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.

By Chencheng Zhu
arXiv Machine Learning
Sep 15

GRADE: Graph Representation of LLM Agent Dependency and Execution

The paper introduces GRADE, a graph-based representation of large language model (LLM) agent executions that captures both execution steps and their dependencies. By adding graded dependency edges—observed, declared, or inferred—to the trace, the authors evaluate how this dependency layer affects failure prediction across six corpora involving tool use, coding, and web tasks. Experiments show that the dependency block can improve prediction in some settings, but its effectiveness varies with the evaluation probe and corpus, and controlled experiments demonstrate that the observed structure is not merely a degree-matched artifact.

By Yue Zhao