Neural swipe decoders are typically tied to the keyboard they were trained on, requiring a new corpus and training run for each layout. In this report, we document our approach toward training models that can function on any contiguous mobile keyboard layout.
arXiv:2606. 31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation.
By Wanxia Cao, Chengzhen Duan, Pei Fu, Pengzhi Gao, Niu Lian, Fazhan Liu, Hui Liu, Heng Qu, Qinzhuo Wu, Zhehao Yu, Tongbo Chen, Shiqi Cui, Anan Du, Shukai Jia, Yuanfa Li, Yike Liu, Wenchao Lu, Haoyuan Sun, Jiatong Sun, Cheng Tan, Yajie Wang, Changqiao Wu, Tao Xiong, Jiahui Yang, Yuxuan Yuan, Ruoceng Zhang, Shaojie Zhang, Jian Zhu, Jian Luan, Cong Zou
arXiv:2605. 18324v2 Announce Type: replace-cross Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders.
By Jaskirat Singh, Boyang Zheng, Zongze Wu, Richard Zhang, Eli Shechtman, Saining Xie
The paper introduces the Sequential Spatio-Temporal Attention Network (SSTAN), a Transformer-based architecture that replaces traditional Graph Convolutional Networks for sign language recognition. SSTAN uses a hierarchical, stacked design with Spatial Multi-Head Attention to model joint relationships within frames and Temporal Multi-Head Attention to capture long-range dependencies across frames, eliminating the need for predefined skeletal graphs. Experiments on large-scale datasets (WLASL, JSL, KSL) show that SSTAN, trained from scratch, achieves state‑of‑the‑art performance in fingerspelling categories and outperforms other skeleton‑only methods on WLASL, highlighting its data efficiency and ability to learn complex spatio‑temporal patterns.
By Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin
arXiv:2607. 00249v1 Announce Type: new Abstract: New device layouts pose a challenging modeling problem due to the lack of large datasets for each specific layout.
By Geeling Chau, Ran Liu, Juri Minxha, Wenhui Cui, Erdrin Azemi, Ellen L. Zippi, Behrooz Mahasseni, Christopher M. Sandino
arXiv:2608. 03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learning are overwhelmingly continuous--joint angles or MANO parameters.
By Haoyu Gu, Haotian Lu, Jingrun Du, Xiao-Ping Zhang
New device layouts pose a challenging modeling problem due to the lack of large datasets for each specific layout. Biosignal foundation models offer a plausible solution if they are able to generalize to new layouts effectively.
arXiv:2601. 18305v2 Announce Type: replace Abstract: Despite numerous Graphical User Interface (GUI) agents claiming to automate user interaction tasks, to date, few achieve satisfactory interaction capability with human users in real-world scenarios.
By Xuan Wang, Siyuan Su, Quantong Fu, Yongxiang Hu, Yangfan Zhou
The paper presents a rotation‑free online handwritten character recognition system that uses Sliding Window Path Signature (SW‑PS) to extract local structural features and a lightweight Linear Recurrent Unit (LRU) classifier. The LRU blends the incremental processing of RNNs with the parallel training efficiency of state‑space models to model dynamic stroke characteristics. Experiments on rotated CASIA‑OLHWDB1.1 subsets (digits, English upper letters, Chinese radicals) achieved accuracies of 99.62%, 96.67%, and 94.33% respectively, outperforming competing models in convergence speed and test accuracy.
By Zhe Ling, Sicheng Yu, Danyu Yang
arXiv:2606. 12629v1 Announce Type: cross Abstract: We show that the standard basis of transformer hidden states already provides a training-free, architecture-general feature basis.
By Varun Reddy Nalagatla
arXiv:2607. 17806v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) requires an embodied agent to interpret a natural-language instruction and predict actions from temporally ordered visual observations.
By Li Xian, Mingxi Li, Yizheng Wang, Yiming Shen, Qi Chen, Zhuoling Xiao
arXiv:2609.23352v1 Announce Type: new
Abstract: Bimanual interaction produces complementary tactile views of the same physical process, yet existing tactile representation learning largely models the...
By Chenxin Liang, Youchen Lai, Chuqiao Lyu, Tianxing Chen, Shoujie Li, Wenbo Ding