What Survives on Real Drawings: Active Sampling, Connectome Wiring, and Matched Baselines in Architectural Document Vision
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.24565v1 Announce Type: new Abstract: Architectural drawings encode material classes through repeated hatch patterns. We test whether a connectome-constrained fly visual network, pretrained...
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid.
The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.
arXiv:2606. 30344v1 Announce Type: cross Abstract: Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression.
The paper studies where to place task‑specific adapters in a vision transformer to balance storage growth and accuracy. Training all contiguous four‑block placements shows an inverted‑U accuracy curve, peaking at intermediate depths, while simple weight or activation metrics favor the deepest blocks. A neuroscience‑inspired method, LS‑B, uses frozen fMRI readouts of human visual areas to select blocks whose responses vary most across tasks, yielding backbone‑specific allocations that match or exceed the best placements found by search and use only 60% of the adapter storage while staying within 1.5 percentage points of full accuracy.
arXiv:2609.31234v1 Announce Type: new Abstract: Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be...