arXiv:2609.05782v1 Announce Type: new
Abstract: Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene a...
By Mohammad Kazzazi, Zixuan Liu, Siavash Khajavi
arXiv:2609.05742v1 Announce Type: cross
Abstract: American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion represen...
By Hongyu Wu, Xu Wu, Tianhao Wu, Jiawei Yu, Phuc Nguyen, Jian Liu, Yi Wu
arXiv:2609.06000v1 Announce Type: cross
Abstract: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet....
By Kiyotaka Kasubuchi, Kazuo Fukiya
arXiv:2609.06663v1 Announce Type: cross
Abstract: Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead...
By Chin Ting Hsu, Yu-Syuan Xu, Ling Zou, Hsien-Kai Kuo, Wen-Huang Cheng
arXiv:2609.08253v1 Announce Type: new
Abstract: Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden...
By Yuehao Wang, Peihao Wang, Hanwen Jiang, Ziyi Yang, Qixing Huang, Zhangyang Wang
arXiv:2509.01809v2 Announce Type: replace-cross
Abstract: We consider the problem of support recovery for sparse binary signals from noisy linear measurements. For sparse Gaussian measurement matrice...
By Youssef Chaabouni, David Gamarnik
The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.
By Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso
The paper explores large‑scale pretraining to enhance deep learning‑based geometric distortion correction for diffusion‑weighted imaging (DWI). By framing the task as image reconstruction, the authors compare a non‑pretrained baseline with self‑supervised and generative pretrained models, finding that the cWDM model yields the best quantitative and qualitative results. When applied to low‑resource, high‑throughput settings in a low‑ and middle‑income country, the pretrained models faced transferability issues, but aligning images to a common standard space improved predictions, indicating that harmonized preprocessing can aid cross‑domain deployment.
By Saroj Khanal, Yashawant Kumar Yadav, Kritam Bhattarai, Jeevan Neupane, Shristi Subedi, Saship Gwachha, Manish Kumar Tiwari, Dong Zhang, Confidence Raymond, Aondona Moses Iorumbur, Udunna Anazodo, Surendra Maharjan, Bishesh Khanal, Mahesh Shakya, Pralhad Kumar Shrestha
arXiv:2609.06974v1 Announce Type: cross
Abstract: Large language models achieve strong performance across diverse tasks, but deployment remains costly because of memory, latency, and energy demands....
By Seungmin Oh, Donggeon Lee, Jongbin Ryu
arXiv:2601.20845v2 Announce Type: replace
Abstract: Time series forecasting is a fundamental problem with applications in climate, energy, healthcare, and finance. Many existing approaches require do...
By Olaf Yunus Laitinen Imanov, Derya Umut Kulali, Taner Yilmaz
arXiv:2609.08368v1 Announce Type: new
Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each st...
By RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su, Andy Ye, Yueming Yuan, Zhichen Zeng
arXiv:2609.09863v1 Announce Type: new
Abstract: Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflict...
By Philip Graemer, Giuseppe Di Caprio
arXiv:2609.08873v1 Announce Type: cross
Abstract: Sparsity is a powerful structural resource in optimization and statistics. We develop frameworks for leveraging sparsity in sampling problems over th...
By Syamantak Kumar, Purnamrita Sarkar, Kevin Tian, Yusong Zhu
arXiv:2609.06898v1 Announce Type: cross
Abstract: Byte Pair Encoding (BPE) constructs vocabularies through greedy pair merging, but the resulting merge order does not necessarily allocate a fixed mod...
By Kenny Shao
arXiv:2609.09300v1 Announce Type: new
Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are dif...
By Zhenxin Qin, Peng Shi, Cong Han, Yinlong Qian, Zequn Jie, Lin Ma
arXiv:2511.02821v2 Announce Type: replace-cross
Abstract: We develop new accelerated first-order algorithms in the Frank-Wolfe (FW) family for minimizing smooth convex functions over compact convex s...
By Dan Garber
arXiv:2601.22475v2 Announce Type: replace
Abstract: Building a generalist robot policy requires continuously integrating new skills while preserving previously acquired behaviors. Directly optimizing...
By Qijun He, Yuxuan Li, Mingqi Yuan, Xiaoquan Sun, Wen-Tse Chen, Jeff Schneider, Jiayu Chen
arXiv:2506.12809v2 Announce Type: replace
Abstract: The long horizon forecasting (LHF) problem has come up in the time series literature for over the last 35 years or so. This review covers aspects o...
By Hans Krupakar, Kandappan V A
arXiv:2609.07557v1 Announce Type: cross
Abstract: 3D Gaussian Splatting has recently revolutionised novel view synthesis as well as many other 3D vision methods and applications. Drawing inspiration...
By Simone Foti, Caner Korkmaz, Stefanos Zafeiriou, Tolga Birdal
arXiv:2609.09924v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational conte...
By Oriol Mar\'in, Roger Mar\'i, Gloria Haro, Rafael Redondo