arXiv:2606. 17961v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer architectures, as it injects information about the spatial or sequential arrangement of inputs.
By Andrea Santomauro, Luigi Portinale, Giorgio Leonardi
arXiv:2609. 20278v1 Announce Type: cross Abstract: Text, knowledge graphs, and hypergraphs all have elements that play distinct roles within relation instances, structure that is lost when data is flattened into token sequences.
By Mahesh Godavarti
arXiv:2607. 10677v1 Announce Type: new Abstract: Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood.
By Binbin Lin, Wei Chen, Yalun Li, Wenxiao Wang, Jieping Ye, Xiaofei He
arXiv:2609.22143v1 Announce Type: cross
Abstract: Machine learning learns functions: prompt to response, image to caption. What these functions are mathematically remains hard to say. We present a me...
By Afjal Chowdhury, James Chen, Alan Edelman
arXiv:2406. 07049v3 Announce Type: replace-cross Abstract: Understanding spatial relationships across all dimensions is fundamental for intelligent systems.
By Boyang Li, Yulin Wu, Nuoxian Huang, Wenjia Zhang
arXiv:2606. 17536v1 Announce Type: cross Abstract: Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry.
By Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang