Brick-Composer: Using MLLMs for Assembly with Diverse Bricks
arXiv:2606. 05445v1 Announce Type: new Abstract: We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks.
arXiv:2606. 07602v1 Announce Type: cross Abstract: LLM-based LEGO assembly generation requires both semantic grounding and physical feasibility.
arXiv:2606. 05445v1 Announce Type: new Abstract: We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks.
arXiv:2503. 19990v4 Announce Type: replace Abstract: Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning across multiple sequential steps.
arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.
arXiv:2608.22637v1 Announce Type: new Abstract: Recent vision-language models (VLMs) show strong capabilities in robotic perception and spatial reasoning, yet their ability to reason about complex me...
Recent vision-language models (VLMs) show strong capabilities in robotic perception and spatial reasoning, yet their ability to reason about complex mechanical assemblies remains underexplored. We int...
The paper introduces a framework that separates physical modeling from execution in physics reasoning tasks. It uses a two‑stage post‑training approach: supervised fine‑tuning to build structured models and reinforcement learning with rubric‑based feedback to refine them. Experiments on PhysReason, PhyX, and SeePhys show that this explicit modeling improves reasoning performance by about 3% on average for small LLMs.
arXiv:2606. 31252v1 Announce Type: new Abstract: Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry.
arXiv:2605. 26182v2 Announce Type: replace Abstract: Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part constraints and structural stability.
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
arXiv:2609.13152v1 Announce Type: new Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experim...
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models (VLMs). By decomposing the problem into planar correspondence (XY), depth consistency (Z), and temporal reversibility (T), FactoSR addresses the dimensional mismatch between 2D visual inputs and 3D physical reasoning. Experiments on multi‑view and video benchmarks show significant performance gains, with a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
arXiv:2510. 14828v3 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully.