arXiv Machine Learning

Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS

arXiv Machine Learning
Aug 4

Real-Time Detection and Repair of LLM Agent Failures

arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.

By Sunny Dubey
arXiv AI
Aug 24

FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning

FL-MAESTRO is a multi‑agent orchestrator that uses three specialized large language model agents to jointly decide the communication topology, per‑client resource allocation, and aggregation rule in each federated learning round. A coordinator merges the agents’ analyses, and a non‑LLM feasibility check validates the decision before execution. By filtering out clients whose updates would never be aggregated, the system eliminates the main source of wasted round energy in volatile edge networks and works across heterogeneous device classes without per‑class energy models, achieving comparable accuracy to the best energy‑aware baseline while reducing wasted energy from over a third to near zero on a non‑IID CIFAR‑10 benchmark.

By Jiajun Wu, Zirui Wang, Jiayu Zhou, Qiang Ye, Steve Drew
arXiv Machine Learning
1d ago

SeedFlood: A Step Toward Scalable Decentralized Fine-Tuning of LLMs

SeedFlood is a novel decentralized fine‑tuning method for large language models that scales to billions of parameters and hundreds of clients. It leverages the seed‑reconstructible structure of zeroth‑order gradients to reduce message sizes to near‑zero, enabling efficient flooding across the network. Experiments show SeedFlood outperforms standard zeroth‑order baselines in communication efficiency and generalization, and rivals first‑order gossip methods while incurring far less communication cost.

By Jihun Kim, Dongyeop Lee, Namhoon Lee
arXiv Machine Learning
Sep 22

When Is Availability-Aware Training Worth It? A Benchmark and Empirical Study of Interruption-Resilient Optimization Under Predictable Compute Schedules

The paper introduces OrbitTrace, a benchmark of 50 physics‑grounded compute‑availability traces from satellite orbits, and investigates whether specialized interruption‑resilient optimizers are needed when training is interrupted by predictable compute gaps. Experiments on CIFAR‑10/ResNet‑18 and GPT‑2/AdamW show that a strong checkpoint‑and‑resume baseline that preserves full optimizer state and indexes learning‑rate schedules in effective time matches uninterrupted training, rendering most availability‑aware methods unnecessary. Only in a narrow regime—large models with non‑persistable optimizer state and frequent short pauses—does reactive adaptation recover a modest portion of the state‑loss penalty, and even this benefit disappears for eclipse‑scale gaps.

By Subhadip Mitra
arXiv Machine Learning
Sep 16

SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation

The paper introduces SWB-DM, a Byzantine‑robust federated learning aggregator that treats each slice of a client update as a one‑dimensional distribution, computes a trimmed Wasserstein barycenter across clients, and uses a medoid‑based gauge‑fixing step to recover coordinate identity. It further incorporates delayed‑momentum caching to decouple robustness from the specific clients sampled each round. Extensive experiments on CIFAR‑10, CIFAR‑100, FEMNIST, and a 500‑client scalability run reveal distinct failure modes of existing defenses and demonstrate that SWB‑DM achieves significant gains, especially when compared under equal round budgets.

By Saranraj S, Saranya M S, Alex David S, Ajay Kumar A