H3DNAS is a hardware‑aware compression framework that operates directly on ONNX computational graphs, eliminating the need for original source code or gradient access. It introduces a Channel Dependency Graph to classify operators and compute a provable compression ceiling, and employs a two‑stage hierarchical search that prunes architectures via L1‑importance channel selection and applies GhostConv mutations to Pareto‑optimal candidates. Applied to 3D point‑cloud models on the ModelNet40 dataset, H3DNAS reduces parameters by up to 65.5% and achieves significant inference speedups with negligible accuracy loss.
By Anchit Mulye, Rhythm Baghel, Sujay Kumar Ingle, Hardik Jain
arXiv:2606. 09924v1 Announce Type: cross Abstract: Deploying deep neural networks on memory-constrained edge accelerators is bottlenecked by per-inference off-chip weight transfer rather than computation: the dense network cannot be retained on-chip, and every parameter must be loaded for every input.
By Kohga Tanaka, Hiroaki Nishi
The paper introduces Neural Spectral Capacity (NSC), a closed‑form metric derived from the singular‑value spectrum of weight matrices that can be computed solely from a network’s architectural specification. Unlike traditional measures such as #Params and #FLOPs, NSC captures architectural structure (depth, width, head, FFN allocations) and can be evaluated without instantiating the model, data, or gradients. Using a dynamic‑programming solver (NSC‑DP), the authors demonstrate that NSC can efficiently identify architectures that outperform existing training‑free proxies across Transformer and CNN families, and achieve state‑of‑the‑art results in tasks such as WikiText‑103 and commonsense reasoning with LLaMA‑7B.
whyItMatters":"NSC provides a fast, architecture‑only proxy that outperforms conventional metrics and training‑free proxies, enabling more effective design and pruning of large models without costly training or data."
By Chenyu Zhu, Ruoyu Zhao, Zhichao Lu
arXiv:2509. 08685v2 Announce Type: replace-cross Abstract: Given encoded 3D point cloud geometry available at the decoder, we study the problem of lossy attribute compression in a multi-resolution B-spline projection framework.
By Tam Thuc Do, Philip A. Chou, Gene Cheung
arXiv:2601. 16622v2 Announce Type: replace-cross Abstract: Equivariant Graph Neural Networks (EGNNs) have become a widely used approach for modeling 3D atomistic systems.
By Lin Huang, Chengxiang Huang, Ziang Wang, Yiyue Du, Chu Wang, Haocheng Lu, Yunyang Li, Xiaoli Liu, Arthur Jiang, Jia Zhang
arXiv:2608. 00859v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients.
By Kazi Ahmed Asif Fuad, Lizhong Chen
arXiv:2602. 00161v2 Announce Type: replace-cross Abstract: In this paper, we formulate the compression of large language models (LLMs) by optimally deleting transformer blocks (``block removal'') as a constrained binary optimization (CBO) problem that can be mapped to a physical system (Ising glass), whose energies are a strong proxy for downstream model performance.
By David Jansen, Roman Rausch, Ali Hashemi, David Montero, Rom\'an Or\'us
arXiv:2607. 06922v1 Announce Type: new Abstract: Deep learning applications have been widely adopted on edge devices, to mitigate the privacy and latency issues of accessing cloud servers.
By Shuo Huai, Di Liu, Hao Kong, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin
The paper introduces GaugeLasso, a method that applies symmetric group‑lasso penalties to transformer channels during training, enabling entire tensor slices to be zeroed out while maintaining dense tensors for GPU efficiency. By calibrating channel penalties based on inference utility per compute, the network self‑organizes into depth‑dependent structural profiles that can be dramatically smaller than the original architecture, achieving up to 255‑fold compression on a polynomial division task and outperforming hand‑designed baselines on language modeling and autoencoding benchmarks. The approach also accelerates training and reveals over‑provisioned axes that guide subsequent design iterations.
By Jed A. Duersch, Na\"im Es-Sebbani, Nathana\"el Haas, Zied Bouraoui
arXiv:2609.36374v1 Announce Type: new
Abstract: Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research gro...
By Brandon Leblanc, Charalambos Poullis
The paper introduces EMR‑HyperNEAT, an eager multi‑resolution grid approach that replaces the recursive quadtree subdivision of ES‑HyperNEAT with a parallelizable evaluation of all grid positions followed by a variance‑based filter. This reformulation removes sequential dependencies, enabling efficient batching across cores and population members, and reduces computational complexity from <O(4^D)> to <O(4^D/P)>. Experiments show 12–34× GPU speedups at depths 5–7 on XOR and higher solve rates across benchmarks.
By Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel
arXiv:2607. 16568v1 Announce Type: new Abstract: Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of the training program.
By Abdallah Khemais (ISITCOM, University of Sousse)