arXiv AI

Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

arXiv Machine Learning
Sep 10

RAPTOR: Role-Aware Private Training for Mixture-of-Experts

arXiv:2609.05770v1 Announce Type: new Abstract: Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see a...

By Duc Dm, Khai Le-Duc, Nguyen Do, Minh Son Hoang, Florent Draye, Thai Hoang, Hoang Phuong Dam, Jiarui Liu, Chris Ngo, Terry Jingchen Zhang, Anh Le Duc Tran, Nhat Do Minh, Minh Ngoc Le, My T. Thai, Ran Xu, Silvio Savarese, Mona Diab, Bernhard Sch\"olkopf, Zhijing Jin, Huy L. Nguyen, Daeyoung Kim
arXiv Machine Learning
Sep 21

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.

By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
arXiv AI
Jul 15

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

arXiv:2607. 12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent.

By Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
arXiv Machine Learning
Sep 16

Multi-Agent Learning with Cooperation-Driven Optimization Dynamics

The paper proposes a cooperation mechanism for multiple small neural network agents that share predictions during training to reduce model complexity while maintaining performance. By incorporating shared predictions into the loss function, agents influence each other's weight updates through strategies such as voter, majority, and weighted average models. Experiments on standard benchmarks show that several small agents can outperform a single large model, achieving comparable accuracy with fewer parameters and lower computational cost.

By Jarod Ketcha Kouakep, Sreyvi UANN, Timoteo Carletti
arXiv Machine Learning
Sep 11

Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based Neuroevolution

The paper investigates a spatial-concentration bias in Evolvable-Substrate HyperNEAT (ES‑HyperNEAT) when applied to MNIST, where evolved networks focus on a central cluster of input pixels. By partitioning the input image into 13 non‑overlapping spatial segments and evolving a separate expert network for each, the authors achieve a 43% mean accuracy—an 106% relative improvement over the baseline—without relying on data‑driven weighting. The study also introduces a receptive‑field diagnostic to detect silent input‑coverage collapse and a spatial‑partitioning remedy to restore full image coverage.

By Romain Claret, Arthur Gygax, Michael O'Neill, Paul Cotofrei, Michael Palma Mendes, Pascal Felber