arXiv AI

Like a Baby: Visually Situated Neural Language Acquisition

arXiv:1805. 11546v3 Announce Type: replace-cross Abstract: We examine the benefits of visual context in training neural language models to perform next-word prediction.

arXiv Machine Learning
Jun 16

Understanding Cross-Modal Contributions in Continual Vision-Language Models: A Theoretical Perspective

arXiv:2606. 14883v1 Announce Type: cross Abstract: Continual vision-language models are commonly addressed through sequential fine-tuning; however, although this paradigm enables adaptation to new environments (tasks), it inherently emphasizes the contribution of previously learned environments (tasks) at the expense of the stability required to preserve previously acquired knowledge.

By Salimeh Sekeh, Mary Wisell