How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
arXiv:2604. 04385v5 Announce Type: replace-cross Abstract: We localize the policy routing mechanism in alignment-trained language models.
Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.
arXiv:2604. 04385v5 Announce Type: replace-cross Abstract: We localize the policy routing mechanism in alignment-trained language models.
arXiv:2606. 29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof.
arXiv:2606. 27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications.
arXiv:2606. 30116v1 Announce Type: new Abstract: Pairwise preference data is widely used for training and evaluating language models (e.
arXiv:2606. 29601v1 Announce Type: cross Abstract: Current languages for specifying multiagent protocols either over-constrain protocol enactments or complicate capturing their meanings.
arXiv:2606. 29548v1 Announce Type: cross Abstract: Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approaching a yellow signal must decide whether to stop or proceed within limited time and distance margins.
arXiv:2606. 22528v2 Announce Type: replace Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget.
arXiv:2606. 30509v1 Announce Type: new Abstract: Matrix factorization (i.
arXiv:2604. 19971v2 Announce Type: replace-cross Abstract: Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking.
arXiv:2606. 29196v1 Announce Type: new Abstract: Do language models know when they are being tested?
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
arXiv:2606. 30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify.
arXiv:2606. 28998v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.
arXiv:2606. 28330v1 Announce Type: cross Abstract: Embedding-based retrieval systems rely on the assumption that geometric proximity in highdimensional representation spaces reflects semantic relevance.
arXiv:2605. 09076v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability.
arXiv:2606. 28410v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) enables text-guided segmentation of unseen objects, breaking fixed-class limitations to achieve open-world understanding.
arXiv:2606. 28416v1 Announce Type: cross Abstract: Deep neural networks (DNNs) have shown outstanding performance in visual recognition tasks within vision sensor networks; however, they are still vulnerable to adversarial manipulations and imperceptible perturbations that can lead to erroneous predictions.
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2606. 28881v1 Announce Type: cross Abstract: Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring systems.
arXiv:2606. 28980v1 Announce Type: cross Abstract: Ovarian cancer is frequently diagnosed at an advanced stage, making preoperative contrast-enhanced computed tomography (CT) central to staging and surgical planning; yet the scarcity of annotated imaging data, compounded by privacy regulations, limits the development of generalizable computational models in this domain.