Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
arXiv:2606. 30989v1 Announce Type: cross Abstract: Warning: This paper contains several toxic and offensive statements.
Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.
arXiv:2606. 30989v1 Announce Type: cross Abstract: Warning: This paper contains several toxic and offensive statements.
arXiv:2606. 30995v1 Announce Type: new Abstract: Recent work has shown that well-optimized individual decision trees can match complex black box models in some settings, primarily in noisy domains.
arXiv:2603. 19127v2 Announce Type: replace Abstract: As Spoken Language Models (SLMs) integrate speech and text modalities, they inherit the safety vulnerabilities of their LLM backbone while introducing an expanded attack surface.
arXiv:2601. 14171v2 Announce Type: replace Abstract: Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between reviewer intent and manuscript details.
arXiv:2606. 31207v1 Announce Type: new Abstract: The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are often sparsely represented in public mobility datasets.
arXiv:2509. 12046v2 Announce Type: replace-cross Abstract: Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned generation remains challenging due to the sparse nature of layout conditions and the risk of feature entanglement.
arXiv:2606. 31976v1 Announce Type: new Abstract: Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domains.
arXiv:2606. 32016v1 Announce Type: new Abstract: Multimodal graph foundation models aim to learn reusable knowledge from graphs enriched with text, images, attributes, and relational topology, thereby supporting diverse graph-centric and modality-centric tasks.
arXiv:2502. 15845v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications.
arXiv:2606. 30935v1 Announce Type: cross Abstract: While neural network control policies are powerful, their deployment on safety critical systems depends on ensuring that they obey strict constraints.
arXiv:2606. 32008v1 Announce Type: new Abstract: Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.
arXiv:2606. 30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric.
arXiv:2606. 30815v1 Announce Type: cross Abstract: Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans.
arXiv:2606. 31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely.
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
arXiv:2606. 30665v1 Announce Type: cross Abstract: Stage B heart failure is characterized by asymptomatic structural or functional cardiac abnormalities.
arXiv:2606. 31876v1 Announce Type: new Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space.
arXiv:2606. 31567v1 Announce Type: cross Abstract: Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety.
arXiv:2603. 26266v3 Announce Type: replace Abstract: Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction.
arXiv:2402. 06734v2 Announce Type: replace-cross Abstract: We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting.