arXiv Computer Vision

A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes

arXiv Computer Vision
Aug 27

A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)

The paper presents a simulator‑grounded framework, 3DTongueQA, that generates verifiable muscle‑grounded question‑answer pairs from 3D tongue meshes. By mapping 11‑dimensional muscle activations to fixed‑topology tongue meshes using the ArtiSynth Badin finite‑element model, the authors produce over 891,000 QA records per language, demonstrating portability across English and Korean. Experiments show that the dataset supports both structured prediction and natural‑language QA, achieving high accuracy with various decoder architectures and robust performance on unseen anchors.

By Seungho Eum, Unsang Park
arXiv AI
Aug 18

Large Language Models and their Awareness of Mechanics and Spatial Geometry

arXiv:2608. 14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically.

By Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber
arXiv Computer Vision
Sep 18

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.

By Hyunjung Chung, Unsang Park
arXiv Computer Vision
Sep 4

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.

By Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden
arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv AI
Aug 19

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

The paper audits frozen decoder‑only large language models (LLMs) on geometric reasoning tasks using parametric CAD constraints. It probes hidden states for linear decodability, forced‑choice generation, activation‑level influence, and behavioral steerability, finding that pretraining improves decoding of local geometric relations but not sketch‑level DOF status. The study shows that decodable information is not always actionable: generation often fails to express it, and steering interventions do not reliably control outputs, revealing divergences among decodability, generation, activation influence, and steerability.

By Man Liang, Xinzhao Cheng, Faizan Wajid
arXiv Machine Learning
Sep 21

Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning

The paper compares two leading motion‑imitation reinforcement learning pipelines—HyFyDy, which uses detailed musculotendon modeling, and MuJoCo, which focuses on computational speed. Using the same human motion‑capture and EMG data, both pipelines reproduce kinematics similarly, but HyFyDy’s muscle activation predictions align more closely with experimental EMG (RMSE 0.164, r = 0.4) than MuJoCo’s (RMSE 0.344, r = 0.11). The authors conclude that HyFyDy’s higher physiological realism makes it currently more suitable for musculoskeletal modeling, though both systems need further development for GPU‑parallelizable environments and robotic assistive‑device design.

By Ayah G. Ahmad, Claire E. Borden, Maegan Tucker
arXiv Machine Learning
Aug 21

Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

arXiv:2608. 19492v1 Announce Type: new Abstract: World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or whether that meaning survives a new action composition.

By Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du
arXiv AI
Jun 8

LuMamba: Latent Unified Mamba for Electrode Topology-Invariant and Efficient EEG Modeling

arXiv:2603. 19100v2 Announce Type: replace Abstract: Electroencephalography (EEG) enables non-invasive monitoring of brain activity across clinical and neurotechnology applications, yet building foundation models for EEG remains challenging due to differing electrode topologies and computational scalability, as Transformer architectures incur quadratic sequence complexity.

By Dana\'e Broustail, Anna Tegon, Thorir Mar Ingolfsson, Yawei Li, Luca Benini
arXiv Computer Vision
Sep 25

AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing

AgenticCADedit introduces a stateful, tool‑mediated approach to multimodal 3D CAD editing, transforming the process from generating a single complete program to executing a sequence of incremental, verifiable actions on a persistent CAD state. By committing each step, inspecting geometry, and selectively reverting faulty operations, the method preserves partial progress and builds upon earlier edits. Experiments across three large language models show substantial gains in validity and acceptance, with the weakest baseline model’s validity rising from 51.0% to 94.8% and a token‑cost reduction of 66.7% compared to neuralCAD‑Edit.

By Saptarshi Neil Sinha, Mika Silvan Goschke, Paul Julius K\"uhn, Arjan Kuijper, Michael Weinmann