arXiv:2607. 14125v1 Announce Type: new Abstract: Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by inserting a class label (e.
By Ruijiang Dong, Zesheng Ye, Jianzhong Qi, Lei Feng, Feng Liu, Gang Niu, Masashi Sugiyama
The paper introduces VISTA, a baseline for dense multi‑label classroom coding that leverages the COPUS protocol as a video benchmark. VISTA applies MiniCPM‑V‑4.5 over sliding windows, refines predictions with an MLP head, and aggregates results onto a 2‑minute COPUS grid, achieving 80.1% macro accuracy on held‑out chemistry lectures. The authors also identify systematic failure modes and provide benchmark tooling and code on GitHub.
By Andrew Franck, Brendan Ng, Ben Fitzgerald, Zane Derrod, Chris Cianci, Chris Craney
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
By Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara, Steven McDonagh
arXiv:2604. 03401v4 Announce Type: replace-cross Abstract: Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns.
By Nolan Platt, Sehrish Nizamani, Alp Tural, Elif Tural, Saad Nizamani, Andrew Katz, Yoonje Lee, Nada Basit
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
By Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
arXiv:2610.02117v1 Announce Type: cross
Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a froze...
By Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris
arXiv:2604. 15336v2 Announce Type: replace-cross Abstract: Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone.
By Shuangquan Feng, Laura Fleig, Ruisen Tu, Philip Chi, Edmund Bu, Melinda Ozel, Junhua Ma, Teng Fei, Virginia R. de Sa
arXiv:2606. 23897v1 Announce Type: cross Abstract: Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predictions on unlabeled domain images.
By Ahmad Algadhi, Ahmed Alzuhair, Omar Alkhulaif, Muzammil Behzad
arXiv:2603.18480v2 Announce Type: replace-cross
Abstract: Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vi...
By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
arXiv:2511. 16107v3 Announce Type: replace-cross Abstract: Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training.
By Shao-Jun Xia, Huixin Zhang, Zhengzhong Tu
arXiv:2607. 18767v1 Announce Type: cross Abstract: The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalability.
By Lachlan McGinness
Prompt learning modifies vision‑language models by optimizing continuous prompt vectors, yet the resulting prompts are hard to interpret in natural language. PromptSpLiCE is a post‑hoc method that rewrites each class‑conditioned text embedding as a sparse mix of concepts from a fixed dictionary, enabling a direct comparison of concept profiles before and after prompt learning. Across 11 image‑classification datasets, the method shows that only about 1.6 of the initial top‑10 concepts remain after learning, and that larger profile changes correlate with higher accuracy gains, while a derived gradient expression offers geometric insight into loss sensitivity.
By Ryo Kamiya, Hiroshi Kera, Kazuhiko Kawamoto