AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,472 stories · RSS feed

arXiv AI
Aug 12

Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

arXiv:2608. 10499v1 Announce Type: cross Abstract: Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy.

By Md Rafid Islam, Rafsan Jany, Zahid Hasan, Ratun Rahman
arXiv AI
Aug 12

TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification

arXiv:2504. 11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual surveys, Bluetooth/WiFi tracking, and Automated Passenger Counters, are often costly, device-dependent, or unable to support individual-level matching.

By Kaicong Huang, Talha Azfar, Jack Reilly, Ruimin Ke
arXiv AI
Aug 12

Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

arXiv:2608. 10273v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs).

By Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong
Hugging Face Trending Papers
Aug 12

A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields

In this paper, we propose a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields. By utilizing the debiased Sinkhorn divergence, our proposed approach develops a differentiable and computationally efficient local distribution matching objective to train stochastic neural networks (SNNs).

Hugging Face Trending Papers
Aug 12

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers.