The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.
By Peidong Wang, Jian Xue, Jinyu Li
The paper introduces a framework for out-of-distribution (OOD) detection that addresses the trade‑off between detection performance and classification accuracy caused by fine‑tuning with auxiliary outlier data. It optimizes three factors—model reminder, data sampling, and representation learning—by proposing Self‑Knowledge Distillation to preserve accuracy, Semi‑hard Outlier Sampling to enhance detection with minimal data, and Outlier‑aware Supervised Contrastive Learning to improve ID‑OOD separability. The combined approach yields cumulative gains, outperforming existing methods on diverse benchmarks, especially in long‑tailed scenarios, and offers a robust baseline for real‑world OOD detection.
By Hyunjun Choi, JaeHo Chung, Hawook Jeong
The paper introduces Uncertainty DMD, a lightweight framework that injects uncertainty into few-step autoregressive video distillation to counteract diversity collapse. By perturbing the first chunk’s timestep and employing a stochastic cache-writing mechanism for subsequent chunks, the method restores stochasticity without altering the model architecture. Experiments demonstrate consistent improvements in video diversity and motion dynamics while preserving visual quality.
By Zixuan Duan, Xunzhi Xiang, Yabo Chen, Xin Zhang, Changhan Liu, Haibin Huang, Chi Zhang, Qi Fan, Xuelong Li
HERALD is a new gradient‑free graph condensation framework that adapts node scoring and feature selection to a graph’s heterophily level. It selects features using a joint Fisher‑discriminability and activation‑density criterion, and scores nodes with a weighted combination of prototype representativeness, decision‑boundary proximity, and Local Intrinsic Dimensionality, where the weights depend on the heterophily ratio. The selected nodes are assembled into a condensed subgraph via score‑ordered BFS expansion, Personalized PageRank pruning, and class rebalancing, achieving comparable storage to BONSAI and outperforming state‑of‑the‑art condensers on heterophilic graphs while remaining competitive on homophilic ones across multiple GNN architectures.
By Sujan Chakraborty, Priyanka Saha, Saptarshi Bej
The paper introduces NOSTRAdAMUS, a predictive link‑adaptation framework for 5G NR that forecasts retransmissions in the next radio frame using recent HARQ history and adjusts the Modulation and Coding Scheme accordingly. Gradient Boosting models achieve 82.9% overall accuracy, with high‑confidence predictions correct 94.2% of the time and a 5.5 µs inference latency. Evaluated OTA on the X5G testbed and various channel emulators, the approach boosts goodput by up to 71.5% and cuts retransmissions by up to 71.8% without retraining across diverse scenarios.
By Tamerlan Aghayev, Maxime Elkael, Michele Polese, Reshma Prasad, Salvatore D'Oro, Yunseong Lee, Koichiro Furueda, Tommaso Melodia
The paper presents a method called Positional Task Conditioning (PTC) to improve defect detection in large product catalogs. By breaking detection into focused sub‑tasks and reinforcing task identity at prompt boundaries, PTC reduces context length and isolates error types, boosting F1 scores from 52% to 87%. The approach outperforms rationale‑based distillation across multiple models, achieving near‑state‑of‑the‑art performance at up to 98% lower cost and is deployed in several countries handling over 10 million product families.
By Soham Satyadharma, Gabriel Roccabruna, Suleiman A. Khan
The paper introduces the Gaussian Belief Propagation Network (GBPN) for depth completion, a hybrid framework that combines deep learning with probabilistic graphical models. GBPN constructs a scene‑specific Markov Random Field via a Graphical Model Construction Network, then infers dense depth distributions using Gaussian Belief Propagation with a serial & parallel message passing scheme. Experiments show GBPN achieves state‑of‑the‑art performance on NYUv2 and KITTI, demonstrating robustness and generalizability across different sparsity levels and patterns.
By Jie Tang, Pingping Xie, Jian Li, Ping Tan
The paper proposes an all‑reflective two‑mirror projection system for EUV lithography that achieves a 4× demagnification at a numerical aperture close to unity (NA≈0.993). Unlike conventional EUV objectives that use 6–10 aspheric mirrors and have <15 % throughput, the design uses a fixed two‑reflection path for each accepted diffraction order, retaining 50–60 % of the power and eliminating order‑dependent phase shifts. The authors optimize 30‑bilayer Bragg coatings for each mirror facet, formulate a 3‑D vector model for a two‑dimensionally periodic mask, and use inverse lithography with a differentiable modal solver to demonstrate simulated sub‑10‑nm aerial images with resolved peaks up to 5 nm defocus.
By Vasiliy A. Es'kin, Egor V. Ivanov, Olga V. Martynova
The paper introduces Weight-Redundancy Pruning (WRP), a forward‑free depth‑pruning technique for large language models that estimates inter‑layer redundancy using only checkpoint weights. WRP compares attention outputs and MLP down‑projection weights across layers, combining pairwise similarities with relative projection‑scale information to guide layer grouping and block selection. Experiments show that WRP consistently outperforms existing forward‑free magnitude pruning methods and approaches the performance of activation‑based pruning across various pruning settings, model families, and downstream tasks.
By Vincent-Daniel Yun, Woosang Lim
The paper investigates how data repetition affects Mixture-of-Experts (MoE) language models compared to dense Transformers. Across models from 80 M to 1 B active parameters, MoEs degrade more quickly as data is repeated, with performance dropping significantly beyond 4× repetition and overtaking dense models only when strong regularization is applied. The study also identifies routing stabilization and expert specialization as key factors in MoE overfitting, and explores regularization techniques that can partially mitigate this issue.
By Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer
The paper introduces a new framework for unsupervised visible‑infrared person re‑identification that leverages modality‑unified prototypes. By contrasting with prototypes that unify both modalities, the method jointly optimizes similarity within and across modalities, improving modality invariance. A self‑distillation step refines instance‑prototype relationships using a steady teacher, resulting in a simple yet effective model validated on standard VI‑ReID benchmarks.
By Menglin Wang, Xiaojin Gong
PitchFlower is a flow‑based neural audio codec that offers explicit pitch controllability by flattening and randomly shifting F0 contours during training while conditioning on the true F0 to reconstruct the original audio. A vector‑quantization bottleneck blocks pitch recovery, and a flow‑based decoder produces high‑quality audio. Experiments demonstrate that PitchFlower matches DSP baselines in pitch accuracy, surpasses state‑of‑the‑art neural codecs in audio quality, and remains robust even when trained on WORLD‑transformed audio, effectively removing vocoder artifacts.
By Diego Torres, Axel Roebel, Nicolas Obin
LILA (Latent-Informed Layer Analysis) introduces a calibration‑free method for structured pruning of large language models by scoring neuron importance using the Kolmogorov–Smirnov distance between singular value distributions of full and neuron‑ablated feed‑forward network weight matrices. The approach requires no training, calibration data, or auxiliary networks, and outperforms existing methods such as PruneNet and SliceGPT on LLaMA‑2‑7B and Phi‑2 at various sparsity levels. After a single epoch of LoRA fine‑tuning, LILA matches heavily calibrated baselines, and a Neural Tangent Kernel analysis provides theoretical support for its spectral importance criterion. Additionally, LILA can dynamically allocate sparsity budgets, achieving state‑of‑the‑art generative preservation and revealing architectural bottlenecks at higher compression.
By Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad
The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.
By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel
SenseNova-U1.5 is an 8B‑MoT native unified multimodal model that can understand, reason about, and generate visual content without using an encoder or VAE. It improves visual fidelity and text rendering through spatially coherent patch reconstruction, large‑scale training on curated generation and editing data, and native resolutions up to 4K. Post‑training, specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing are optimized and distilled into a multi‑expert framework, yielding advances in image fidelity, complex composition, multi‑reference editing, and instruction following.
By Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng, Jiangnan Chen, Ruixi Zhang, Ruohui Wang, Wenwen Tong, Xiangyu Fan, Yubo Wang, Yue Zhu, Yuwei Niu, Zhengqi Bai, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Bo Yang, Chen Feng, Chengguang Lv, Guangjia Liu, Guanlin Wang, Hanyu Zhang, Haojia Yu, Hongcan Xiao, Hongli Wang, Huan Wu, Huaping Zhong, Jian Fang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jing Zuo, Jingcheng Ni, Junxiang Xu, Linjun Dai, Mutian Xu, Peishen Yan, Penghao Wu, Ruijie Mao, Ruisi Wang, Shihao Bai, Shuang Yang, Shuya Yang, Shuyan Zheng, Silei Wu, Siying Li, Tao Chu, Tianbo Zhong, Tongxi Zhou, Weichao Luo, Weichen Fan, Wenhao Jia, Wenjie Gao, Xiangli Kong, Yan Li, Yang Yong, Zimo Wen, Zixuan Qian, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin
The paper introduces SEM‑HD, a framework that leverages longitudinal mammography history as privileged information during training to improve risk prediction while requiring only a single current exam at inference. By having a student model predict latent representations of past visits and using teacher supervision from actual longitudinal data, SEM‑HD preserves temporal modeling benefits without needing prior exams at deployment. Experiments on three cohorts and two backbone architectures show consistent gains in long‑horizon AUC and pAUC, especially in low false‑positive‑rate regions, and recover much of the performance gap to full‑history models.
By Banafsheh Karimian, Soufiane Belharbi, Alexis Guichemerre, Luke McCaffrey, Mohammadhadi Shateri, Eric Granger
BiHDTrans is a neurosymbolic binary hyperdimensional transformer that merges self‑attention with hyperdimensional computing to classify multivariate time series efficiently. It surpasses existing HD models by at least 14.47% and binary transformers by 6.67% on average, while an FPGA‑accelerated implementation reduces inference latency 39.4× compared to state‑of‑the‑art binary transformers. Even with a 64% reduction in hyperspace dimensionality, BiHDTrans remains competitive, achieving 1–2% higher accuracy with 4.4× smaller model size and nearly 50% lower latency than the full‑dimensional baseline.
By Jingtao Zhang, Yi Liu, Qi Shen, Changhong Wang
The paper investigates associative memory in a bipartite Hopfield–Krotov architecture, termed class H, where hidden neurons serve as the retrieval order parameter. Using the replica method, it derives replica‑symmetric phase diagrams and closed‑form capacities for polynomial load, showing that crosstalk statistics are similar for Ising and spherical visible neurons. With a softmax hidden layer, the load becomes exponential, mapping the thermodynamics onto a random‑energy‑model that exhibits paramagnetic, condensed, and frozen phases, and revealing that heating destabilizes retrieval through quantized attention reassignments while Gaussian patterns remain metastable at all loads.
By Toshihiro Ota, Masato Taki
HiRAD is a hierarchical reinforcement learning framework designed for continuous-space routing of large-scale AGV fleets, offering real-time guarantees. It introduces a step-level spatiotemporal representation, separates heading selection from velocity control to shrink the action space, and employs an asynchronous event-driven decision pipeline that reduces inference complexity from O(n²) to O(n) and cuts per-step latency by up to 71%. Experiments on random graphs and two warehouse maps show that HiRAD decreases makespan by 45% to 63% and shortens overall runtime.
By Yunjie Huang, Ruizhong Wu, Mengxuan Zhang, Frodo Kin Sun Chan, Yan Nei Law, Lei Li
The paper introduces a method that integrates a frozen text‑to‑image diffusion model into density‑based topology optimization via score distillation sampling. By converting a natural language prompt into a generative gradient, the approach lets physics decide which design features survive, achieving significant compliance reductions across multiple domains and physics regimes. An automated pipeline then transforms the optimized density fields into CAD‑ready geometries.
By Yongmin Kwon, Namwoo Kang