Voxtral
Related stories
Speaking of Voxtral
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.
DVD: Discrete Voxel Diffusion for 3D Generation and Editing
arXiv:2605. 07971v2 Announce Type: replace-cross Abstract: We introduce Discrete Voxel Diffusion (DVD), a discrete diffusion framework to generate, assess, and edit sparse voxels for SLat (Structured LATent) based 3D generative pipelines.
Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms
Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, including autonomous navigation and augmented reality. In this study, a systematic review has been conducted to compile the research contributions made in the last ten years, i.
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
arXiv:2607. 12752v1 Announce Type: cross Abstract: While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry.
Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms
arXiv:2606. 24180v1 Announce Type: cross Abstract: Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, including autonomous navigation and augmented reality.
HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation
arXiv:2607. 13468v1 Announce Type: cross Abstract: Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation.
FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native fMRI
arXiv:2606. 11500v1 Announce Type: cross Abstract: The success of large-scale deep learning models in neuroscience is fundamentally constrained by severe data heterogeneity.
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
arXiv:2510. 05356v2 Announce Type: replace-cross Abstract: Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score function, which in turn leads to interpolations between modes of the data distribution.
The Hallucinations Leaderboard, an Open Effort to Measure Hallucinations in Large Language Models
P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization
arXiv:2608. 07549v1 Announce Type: cross Abstract: Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem: how to sample and organize geometric evidence into compact, structured and learnable tokens.
Discovering Functionally Selective Brain Regions with a Deep Topographic Multimodal Model
arXiv:2606. 09770v1 Announce Type: cross Abstract: Nearby neurons in cortex share similar response profiles, producing systematic spatial organization across sensory and cognitive systems.