AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,629 stories · RSS feed

arXiv AI
Aug 5

CollaFuse: Collaborative Diffusion Models

arXiv:2406. 14429v4 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic images.

By Simeon Allmendinger, Domenique Zipperling, Lukas Struppek, Niklas K\"uhl
arXiv Machine Learning
Aug 5

UNVaMP: Neural Knowledge Tracing with Variational Regularization of Latent Knowledge Dynamics

arXiv:2608. 03811v1 Announce Type: new Abstract: We introduce the Unified Neural Variational Measurement of Proficiency (UNVaMP) architecture, a knowledge tracing method that integrates observed student-item interactions with internal memory to produce evolving latent representations of student knowledge.

By Carson J. Cook, Ahmed J. Zerouali, Anthony Schmidt, Reginald Ziedzor, Paul Lin, Luke G. Eglington