The paper demonstrates that a preference‑optimization objective can learn to distinguish reliable from unreliable sources by installing a prior‑dependent reliability switch. By training on data where a source’s stated reliability is paired with its answer, the model learns to flip its response only when the stated reliability exceeds a threshold that grows with the model’s prior. Experiments on Qwen2.5‑7B‑Instruct and Llama‑3.1‑8B show that this switch generalizes to unseen reliability values and follows stated reliability over role prestige, whereas supervised imitation fails to learn it.
By Sen Yang, Yuen-Hei Yeung
arXiv:2607. 02104v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2607. 05806v1 Announce Type: new Abstract: Training data for machine learning is routinely collected by a selection process the model never sees: loans are observed only when granted, outcomes only when a test was ordered.
By Gunner Levi Howe
The paper investigates how an informed adversary can influence the optimal signal in a constrained signalling channel. It finds that the adversary‑robust optimum aligns with the salience pole on most items, differing only on a small subset where the salience‑to‑Bayes coordinate is undefined. The study demonstrates that as the adversary’s persuasion budget increases, the optimal signal shifts from a posterior‑maximizing to a margin‑maximizing strategy, and provides a diagnostic check for evaluating adversary‑awareness.
By Cris Huynh
arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.
By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
The paper studies contextual bilateral trade with full feedback, showing that action-independent observations eliminate the usual polynomial adaptation penalty seen in heavy-tailed bandits. It presents fully parameter-free algorithms that achieve oracle minimax regret rates without knowing the moment order or scale, and derives new regret bounds for both parametric and nonparametric settings. The key technical insight is a paired squared‑loss statistic whose noise cancels, enabling model selection and yielding regret rates that interpolate between classical nonparametric and linear extremes.
By Hangyi Zhao