arXiv Machine Learning By Yaohui Cai, Vesal Bakhtazad, Cunxi Yu, Zhiru Zhang

GauS: Differentiable Scheduling Optimization via Gaussian Reparameterization

Read the original on arXiv Machine Learning →

arXiv:2602. 20427v2 Announce Type: replace Abstract: Efficient operator scheduling is a fundamental challenge in software compilation and hardware synthesis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 15

Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution

The paper introduces a partition-aware scheduling framework for mobile inference on heterogeneous platforms that combines mobile GPUs and multiple CPU core clusters. It jointly optimizes operator partitioning, device assignment, and execution order for static DAGs of operators, such as those in CNNs or vision transformers. An online iterative search approach decomposes large DAGs into stages, targets critical operators, and uses latency predictors to avoid exhaustive profiling, achieving near‑optimal latency with minimal scheduling overhead.

By Zhuojin Li, Marco Paolieri, Leana Golubchik
arXiv AI
Aug 20

Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions

The paper introduces SDDL, a neuro‑symbolic framework that converts natural‑language combinatorial scheduling problems into compact, solver‑aligned representations, delegating low‑level modeling and search to a deterministic compiler and external solver. On a 300‑instance subset of scheduling tasks, SDDL achieves higher feasibility rates for resource‑constrained language models—up to 55.3% and 28.3%—compared to direct‑generation baselines (23.7% and 1.3%) and solver‑code baselines (21.7% and 7.0%), with a median optimality gap of 0.0% among feasible schedules.

By Shrenil Shaun Sharma, Avi Sharma
arXiv AI
Sep 3

MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs

The paper introduces a MeanField surrogate model for predicting performance of concurrent heterogeneous AI inference workloads on shared GPUs, reducing profiling complexity from combinatorial to linear in the number of models. Experiments with up to six models show high accuracy (R²≈0.96) and efficient integration into a genetic algorithm scheduler, achieving near-exhaustive search performance with minimal runtime overhead.

By Youssef Ennouri, Soonhoi Ha