Valley3: Scaling Omni Foundation Models for E-commerce
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
TLive-Omni is an omni‑modal understanding model designed for e‑commerce live streaming, integrating image, video, audio, and text inputs into a unified representation. It introduces Per‑vGrid for timestamped token organization, a three‑stage supervised training pipeline, and a Faithful‑RFT reinforcement fine‑tuning stage to enhance answer faithfulness and expression quality. The model is supported by a scenario‑oriented capability taxonomy and a compact data production engine that generates training signals for tasks such as speech recognition, product visual grounding, and omni‑modal QA, achieving strong performance on live‑commerce benchmarks and good generalization to general tasks.
arXiv:2602. 22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world.
arXiv:2609.18323v1 Announce Type: new Abstract: Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H...
arXiv:2606. 15231v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex, open-world scenarios.
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.