arXiv AI By Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

Read the original on arXiv AI →

arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.