arXiv:2609.07557v1 Announce Type: cross
Abstract: 3D Gaussian Splatting has recently revolutionised novel view synthesis as well as many other 3D vision methods and applications. Drawing inspiration...
By Simone Foti, Caner Korkmaz, Stefanos Zafeiriou, Tolga Birdal
arXiv:2608. 16080v1 Announce Type: new Abstract: Thermal-aware optimization of multi-die 3D integrated circuits evaluates many designs, each a costly heat-equation solve.
By Xinling Yu, Yixing Li, Ziyue Liu, Xin Ai, Zhiyu Zeng, Hai Li, Zheng Zhang
arXiv:2609.36692v1 Announce Type: cross
Abstract: Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram represen...
By Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu
arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.
By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang
The paper introduces Spectrally Optimised Neural Discretisations (SpeND), a mesh‑free framework that learns discretisation weights from local stencil geometry on unstructured point clouds. By embedding discrete moment conditions into the network architecture, SpeND guarantees polynomial consistency and allows the weights to be optimised for spectral accuracy over a chosen wavenumber band, using an unsupervised Fourier‑mode loss. The resulting operators are PDE‑agnostic, perform well on Poisson, Burgers, and Navier–Stokes equations, and can reduce wall‑clock time by 3–20× compared to existing mesh‑free methods at the same accuracy.
By Lucas Gerken Starepravo, Henry Broadley, Steven Lind, Jack R. C. King
arXiv:2606. 26453v1 Announce Type: new Abstract: We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating large language model (LLM) code generation with hardware profiler feedback and pluggable bottleneck detection tools.
By Jiading Gai, Shuai Zhang, Kaj Bostrom, Jin Huang, Vihang Patil, Haoyang Fang, Bernie Wang, Huzefa Rangwala, George Karypis
arXiv:2607. 03949v1 Announce Type: cross Abstract: Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings.
By Zhengpeng Feng, Sadiq Jaffer, Ira Shokar, Jovana Knezevic, Mark Elvers, Clement Atzberger, Robin Young, Aneesh Naik, Niall Robinson, Andrew Blake, David Coomes, Anil Madhavapeddy, Srinivasan Keshav
arXiv:2606. 17603v1 Announce Type: new Abstract: In Self-Supervised Learning (SSL), preventing representation collapse by explicitly enforcing a uniform distribution on the unit hypersphere has proven to be effective.
By L\'eo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Max Dunitz (ATT), Marc Pic (ATT), Pablo Mus\'e (CB, IFUMI), Gabriele Facciolo (CB)
Heat Field Signatures (HFS) lift irregular point clouds into a multiscale family of smooth ambient heat fields, enabling closed‑form computation of global and local geometric signatures directly from pairwise distances. HFS captures heat concentration, intrinsic dimension, anisotropy, and scale transitions, and introduces the Heat Dimension Spectrum (HDS) as a compact multiscale summary. The method serves as a descriptor, lightweight learned representation, or feature channel for neural point‑cloud models, outperforming strong baselines on synthetic and real‑world benchmarks while reducing end‑to‑end cost.
By Yuanqing Wang, Yapeng Tian, Baris Coskunuzer
arXiv:2608. 05033v1 Announce Type: cross Abstract: Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning.
By Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding
arXiv:2607. 14568v1 Announce Type: cross Abstract: A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.
By A. C. Opus, J. Q. Lu
Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e. scipy.ndimage, is CPU-only, single-array, and th...