Test-Time Detoxification without Training or Learning Anything
arXiv:2602. 02498v2 Announce Type: replace-cross Abstract: Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale.
Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.
arXiv:2602. 02498v2 Announce Type: replace-cross Abstract: Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale.
arXiv:2601. 22823v2 Announce Type: replace-cross Abstract: We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions.
arXiv:2510. 18989v2 Announce Type: replace Abstract: Neural operators are commonly utilized as fast surrogates for numerical solvers in PDE problems, mapping input functions to solution functions.
arXiv:2606. 29581v1 Announce Type: cross Abstract: Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluations usually treat these choices as fixed implementation details.
arXiv:2606. 23129v2 Announce Type: replace-cross Abstract: Implicit Neural Representations (INRs) have been proven successful in encoding continuous signals through coordinate-based networks, yet facing a spectral dilemma: periodic activations capture fine details but act as all-pass filters that memorise noise, while spatially compact activations regularise effectively but suffer from low-frequency bias.
arXiv:2606. 25178v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science.
arXiv:2605. 14252v2 Announce Type: replace-cross Abstract: Spiking neural networks (SNNs), which are brain-inspired and spike-driven, achieve high energy efficiency.
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
arXiv:2606. 29240v1 Announce Type: new Abstract: Heterogeneous graph neural networks (HGNNs) have achieved strong performance in modeling complex graph-structured data with multiple node and relation types.
arXiv:2606. 29196v1 Announce Type: new Abstract: Do language models know when they are being tested?
arXiv:2606. 29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
arXiv:2601. 03546v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions.
arXiv:2504. 10796v4 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss can lead to overly conservative policies.
arXiv:2510. 25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits.
arXiv:2606. 28963v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated.
arXiv:2606. 30587v1 Announce Type: cross Abstract: Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection.
arXiv:2602. 13562v2 Announce Type: replace-cross Abstract: While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures.
arXiv:2606. 29867v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversarial perturbations that can severely degrade performance.
arXiv:2606. 28716v1 Announce Type: new Abstract: The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems.
arXiv:2606. 29322v1 Announce Type: new Abstract: Collaborative learning is sustainable only when it benefits each participant.