Verifiably grounded machine interpretation of lunar geology
arXiv:2608. 09276v1 Announce Type: cross Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations.
arXiv:2606. 25000v1 Announce Type: new Abstract: To evaluate whether vision-language models can reason about geological histories, it is necessary to construct observations for which the underlying process history is known.
arXiv:2608. 09276v1 Announce Type: cross Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations.
arXiv:2605. 03383v2 Announce Type: replace Abstract: Geological interpretation infers subsurface properties and structures from indirect geophysical observations.
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models (VLMs). By decomposing the problem into planar correspondence (XY), depth consistency (Z), and temporal reversibility (T), FactoSR addresses the dimensional mismatch between 2D visual inputs and 3D physical reasoning. Experiments on multi‑view and video benchmarks show significant performance gains, with a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
CoEvolve is a construct-to-edit framework for visual grounding that separates the task into explicit state construction and state editing. It uses Region‑Evolution Reinforcement to progressively refine candidate regions and Bidirectional Denoising Refiner to adjust coordinate fields based on fixed semantic context. The approach achieves high grounding accuracy with a 9B backbone, rivaling much larger models, and can recover over 27 percentage points in mean box overlap after a single refinement pass.
arXiv:2508. 07683v2 Announce Type: replace-cross Abstract: Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries.
arXiv:2609.20978v2 Announce Type: replace Abstract: High-consequence subsurface decisions often rely on sparse data that permit competing geological interpretations. Determining consistency of these...
The paper introduces V‑Rubrics, a reinforcement‑learning framework that evaluates vision‑language model responses by breaking them into atomic propositions and scoring them on Visual Faithfulness, Reasoning Consistency, and Instruction Following. Using a fine‑tuned Qwen3‑VL‑8B‑Instruct model and a newly created 50K‑example V‑Rubrics dataset, the authors demonstrate that rubric‑based GRPO outperforms both a shared SFT baseline and an answer‑only GRPO, especially on knowledge‑oriented and visually grounded reasoning tasks.
arXiv:2511.14086v2 Announce Type: replace-cross Abstract: Despite recent progress in 3D-LLMs, they remain limited in accurately grounding language to visual and spatial elements in 3D environments. T...
arXiv:2606. 24967v1 Announce Type: new Abstract: In ill-posed inverse problems, the recovered solution depends as much on the prior as on the data, yet much of the engineering knowledge that could serve as that prior is recorded qualitatively rather than in formal mathematical form.
arXiv:2608. 02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains.
arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.