arXiv AI

Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation

arXiv:2605. 28642v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT).

arXiv Machine Learning
5d ago

LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant

LUMO (Lightweight Unified Multilingual Orchestrator) is a privacy‑preserving offline voice assistant that runs entirely on edge hardware, specifically a Raspberry Pi 5 with 8 GB RAM. It integrates local ASR, a 4‑bit GGUF‑quantized LLM, and TTS to deliver end‑to‑end response latencies of 2.0–4.0 s, a 6.8 % WER on short English utterances, and lower peak power consumption (~9 W) compared to existing edge assistants. The system also supports Bangla speech, enabling multilingual use in low‑resource settings.

By Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
arXiv AI
Sep 21

Samsone: A Family of Open Small Audio Language Models for On-Device Inference

Samsone is a family of small audio language models (SALMs) designed for on‑device inference, with the flagship Samsone‑134M setting a new state‑of‑the‑art for its size class across multiple benchmarks. The paper also presents Samsone‑99M and Samsone‑356M to study scaling laws, showing that these compact models achieve performance competitive with much larger counterparts. The authors train the models on publicly available data and release training code, weights, mobile‑optimized checkpoints, and an open‑source Android app for real‑time inference.

By Piotr Masztalski, Micha{\l} K. Grzeszczyk, Olaf Sikorski
arXiv Machine Learning
Sep 11

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

The paper introduces EMMI, a framework that enables communication‑efficient inference of multimodal large language models (MLLMs) on edge devices. EMMI encodes each sensor modality separately, fuses the representations, and compresses them into a compact latent vector that is transmitted to a server for high‑capacity reasoning. Experiments on a multimodal benchmark show that EMMI can cut the communication payload by 32× while keeping accuracy comparable, achieving up to a 3.4× reduction in end‑to‑end inference latency under bandwidth‑constrained conditions.

By Motahare Mounesan, Irfan Khan
arXiv AI
Sep 7

Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

Diffusion language models (DLMs) provide a non‑autoregressive approach for mobile edge agentic AI, refining tokens through iterative denoising instead of left‑to‑right decoding. They can update multiple uncertain tokens in parallel and use bidirectional context, allowing flexible quality‑latency trade‑offs and early exits that reduce response delay and communication overhead. The survey reviews DLM foundations, resource‑efficient architectures, training and inference acceleration, compression, deployment strategies, and discusses open issues such as long‑context management, split inference, and trustworthy execution.

By Chenqi Li, Minghui Min, Dusit Niyato, Wei Ni