Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling.
arXiv:2606. 16327v1 Announce Type: cross Abstract: Recent acoustic-to-articulatory inversion (AAI) models rely on electromagnetic articulography (EMA) data, which are costly and limited in scale.
By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv:2609.09757v1 Announce Type: cross
Abstract: Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, gra...
By Hong Nguyen, Sean Foley, Christina Hagedorn, Yijing Lu, Sudarsana Reddy Kadiri, Dani Byrd, Shrikanth Narayanan
The paper introduces a method for measuring accent differences that balances interpretability and practicality. It proposes using articulatory representations obtained via articulatory inversion as an interpretable basis for accent comparison, while employing optimal transport to compare accents across any type of recording. This approach aims to overcome the limitations of traditional phonetic analyses and embedding‑based methods, which are either time‑consuming or non‑interpretable.
By Charles McGhee, Mark J. F. Gales, Kate M. Knill
arXiv:2609.36737v1 Announce Type: cross
Abstract: The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable an...
By Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann
KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.
By Hyunjung Chung, Unsang Park
arXiv:2607. 10551v1 Announce Type: cross Abstract: Accurate geometric calibration is essential for fluoroscopy-guided spinal imaging, digitally reconstructed radiograph (DRR) generation, and 2D--3D vertebral registration.
By Lin Li, Chaochao Zhou, Benjamin Aubert, Junlin Guo, Junchao Zhu
The paper proposes using 3D Gaussian Splatting to reconstruct volumetric MRI data from non‑aligned scans, enabling the generation of novel view planes that are better aligned with spinal anatomy. These resampled images are then used to predict ordinal severity grades of localized stenosis, outperforming traditional Voxel Interpolation and Cubic B‑spline resampling methods. Across all stenosis conditions, the Gaussian Splatting approach yields more accurate diagnostic gradings than raw or conventionally resampled scans.
By Robin Y. Park, Mark C. Eid, Rhydian Windsor, Amir Jamaludin, Ana I. L. Namburete, Jo\~ao F. Henriques, Andrew Zisserman
arXiv:2607. 19696v1 Announce Type: cross Abstract: The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated systems are hindered by the lack of diverse, high-quality benchmarks.
By Duong Ngoc Vu, Hai Son Nguyen, Trong-Nghia Nguyen, Bien Tran Van, Trang Mai Xuan, Huan Vu, Thien Van Luong
arXiv:2609.24560v1 Announce Type: new
Abstract: Confocal Laser Endomicroscopy (CLE) provides real-time, cellular-resolution optical biopsy but has a narrow field of view, which image mosaicing can ex...
By Ahmed Aboelela, Johannes Barcsay, Jana Friedhof, Miguel Gon\c{c}alves, Alexander Hann, Katharina Breininger
arXiv:2606. 15250v1 Announce Type: cross Abstract: Radiographic assessment of lower-limb alignment (LLA) is important for predicting joint health and surgical outcomes in total knee arthroplasty.
By Zhisen Hu, Antti Kemppainen, David Johnson, Egor Panfilov, Huy Hoang Nguyen, Timothy Cootes, Claudia Lindner, Aleksei Tiulpin
arXiv:2609.38019v1 Announce Type: new
Abstract: We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rat...
By Bangxun Tang