arXiv Machine Learning

Closing the Social-Semantic Gap: SPSD for Edge-Based Prompt Compression in Cloud LLM Inference

arXiv:2606. 19364v1 Announce Type: new Abstract: The prefill stage of Large Language Model (LLM) inference is a growing contributor to cloud-scale energy cost.

arXiv Computation and Language
Sep 4

ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

ParaBridge is a self‑distillation method that uses a temporary paralinguistic instruction scaffold during training to teach a speech language model when non‑lexical cues should influence dialogue responses. By providing dense, full‑vocabulary next‑token targets from the scaffolded view while the scaffold‑free model generates its own replies, ParaBridge stabilizes inference‑time behavior without requiring curated dialogues or external reward models. Experiments on Qwen3‑Omni show significant gains on safety and empathy benchmarks while preserving general performance across multiple tests.

By Yuxiang Wang, Qinke Ni, Shengbo Cai, Wan Lin, Liqiang Zhang, Zhizheng Wu
arXiv AI
Sep 17

When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI

The paper evaluates three approaches for emotion recognition in conversation— a low‑cost stacked ensemble, an off‑the‑shelf LLM prompt, and a confidence‑gated hybrid that escalates only uncertain ensemble predictions to the LLM. Across three datasets (IEMOCAP, MELD, CMU‑MOSI), the hybrid consistently outperforms each pure system, achieving higher weighted F1 scores while routing most traffic through the inexpensive ensemble. This results in significant cost savings (≈$10‑85 per million utterances) and provides an interpretable escalation signal tied to emotion or sentiment shifts.

By Sai Babu Udayagiri, Arjun Chouhan, Ravisekhar Kanagala, Trishala Pavagada
arXiv AI
Sep 2

Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs

The paper compares the sustainability of native small language models (SLMs) versus large language models (LLMs) compressed via post‑training quantization for edge AI deployment. Using a Holistic Sustainability Score (HSS) that balances capability, efficiency, and safety, the study evaluates 30 configurations across five benchmarks, latency, VRAM, energy, and harmful‑prompt robustness. Results show that optimized quantized LLMs can outperform SLMs overall, while SLMs remain competitive due to lower resource demands, challenging the assumption that native SLMs are always the most sustainable choice.

By Jainil Dharmil Shah
arXiv Machine Learning
Aug 28

A Survey of LLM Prompt Datasets: Taxonomy, Linguistic Patterns, and Practical Uses

The paper presents a survey of 129 public large language model (LLM) prompt datasets, totaling over 1.22 TB and 673 million instances, and introduces a unified taxonomy for them. By analyzing seven datasets in depth, the authors identify lexical, syntactic, and semantic patterns that differentiate prompts from general text, and evaluate these patterns for tasks such as prompt filtering, source domain routing, and response quality assessment. They demonstrate that a 63‑dimensional linguistic feature set extracted on a CPU can match over 91 % of the F1 score of GPU‑based sentence embeddings while halving latency, and that structural features can effectively route prompts across datasets, though they may negatively impact response quality when prompt length is controlled.

By Yuanming Zhang, Yan Lin, Arijit Khan, Huaiyu Wan