arXiv Machine Learning

Towards a future space-based, highly scalable AI infrastructure system design

arXiv:2511. 19468v2 Announce Type: replace-cross Abstract: If AI is a foundational general-purpose technology, we should anticipate that demand for AI compute -- and energy -- will continue to grow.

arXiv AI
Sep 15

Deep Tech to Space: Space Data Centers and AI Revolution at the Edge

arXiv:2605.19892v2 Announce Type: replace-cross Abstract: Dramatic cost reductions driven by private sector innovations have led to a rapid increase in the number of satellites in orbit and a corresp...

By Jonas Weiss, Patricia Sagmeister, Gabriel Maiolini Capez, Dinesh Verma, Roberto Garello, Alberto Perotti, Dawid Lazaj, Alicja Musial, Jakub Nalepa, Thomas Morf, Martin Schmatz, Marek Krawczyk, Mateusz Przeliorz, Kevin Roche, Sagar Tayal, Mahalakshmi Lakshminarayanan, Nicolas Long\'ep\'e, Pierre-Philippe Mathieu, Agata Wijata
arXiv AI
Aug 26

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

ShardMeter is a lightweight analytical performance model that predicts end-to-end runtime for transformer-based workloads across sharded, distributed, and decentralized training setups. By taking a model’s characteristics and a target hardware topology as input, it estimates per-GPU and per-island throughput, training cost, total wall-clock time, and pinpoints performance bottlenecks. The model reveals diminishing-return regimes with increasing island size, quantifies compute- versus communication-bound scaling, evaluates hyperparameter trade-offs, and models cost-throughput for large-scale decentralized training, enabling rapid exploration of configuration space and near-optimal deployment plans.

By Tim Beringer (Technical University of Darmstadt), Patrick Diem (Technical University of Darmstadt), Felix Wolf (Technical University of Darmstadt), Arya Mazaheri (Technical University of Darmstadt, PanocularAI)
arXiv AI
Aug 26

ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal

The paper introduces OrbitALIF, a federated learning framework that performs cloud removal on low‑earth‑orbit satellites. It uses a compact 2.30 M‑parameter spiking neural network with adaptive gated fusion and spectral‑spatial hybrid attention modules, enabling both training and inference onboard. The approach achieves competitive cloud‑removal quality while consuming only 0.287 mJ per inference on neuromorphic hardware, a 72.3‑fold energy reduction compared to an equivalent ANN.

By Bohan Zhang, Chenyu Xu, Yijie Mao, Yuanming Shi
arXiv Machine Learning
Jul 13

A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

arXiv:2607. 09084v1 Announce Type: new Abstract: The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability.

By Linhui Xiao, Guiping Cao, Mingyue Guo, Xianchao Guan, Fan Yang, Ming Tao, Xin Li, Yuxin Peng, Yaowei Wang
arXiv AI
Sep 2

FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study

The paper introduces a heterogeneous federated learning approach using the FractalNet architecture tailored for satellite mega‑constellations. It formalizes contact‑window‑constrained, depth‑heterogeneous optimization and proposes a distributed path scheduler that assigns model depth based on satellite SWAP‑C constraints, predicted contacts, and training statistics. The framework includes periodic update pooling and a three‑tier agentic control plane, and is validated through a wildfire detection case study across LEO, MEO, and GEO/HEO shells, demonstrating improvements in convergence, communication efficiency, energy adaptation, and robustness.

By Sai Puppala, Koushik Sinha
arXiv Machine Learning
Sep 18

Traffic Engineering in Large-scale Networks with Generalizable Graph Neural Networks

The paper introduces TELGEN, a traffic engineering algorithm that uses graph neural networks to predict an optimal TE algorithm rather than a direct solution. TELGEN generalizes across diverse network topologies and traffic patterns, achieving less than a 3% optimality gap on networks up to 5,000 nodes and 3.6 million links, while reducing solving time by up to 84% and training time by up to 79.6% compared to existing methods.

By Fangtong Zhou, Xiaorui Liu, Ruozhou Yu, Guoliang Xue
arXiv Machine Learning
Sep 22

Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing

The paper introduces the Universal Observatory Graph (UOG), an AI‑driven framework that models autonomous observatories at L2 Lagrange points as nodes in a weighted graph, with edges defined by interplanetary distance, latency, transmission power, and reliability. Using a six‑observatory Solar System configuration (Earth, Mars, Jupiter, Saturn, Uranus, Neptune), the authors evaluate instantaneous sky coverage with three methods, all showing complete network union coverage and modest overlap. They formulate communication routing as a finite‑horizon Markov decision process solved via tabular Q‑learning, identifying the Earth‑Saturn‑Uranus‑Neptune path as the highest‑return route among 41 feasible simple paths under a four‑hop constraint.

By Mohammed Abdel Razek