arXiv:2608.08477v5 Announce Type: replace
Abstract: We build VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model for offline use, and measure what limits its visual groundi...
By Juan S. Santillana
arXiv:2609.38956v1 Announce Type: new
Abstract: Routing signals of modern vision transformers -- expert gates, attention-residual weights and halting scores -- often improve probes that predict wheth...
By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv:2608. 04084v1 Announce Type: new Abstract: Mixture-of-experts (MoE) networks pursue specialization through learned routers, gates, and load-balancing losses, yet at matched total-parameter budgets learned routers can underperform equal-weight No-Routing baselines.
By Boyao Wang, Zhihan Lei
The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.
By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon evaluation framework for vision‑language models that tests hidden‑target identification across galleries of visually similar images. The benchmark requires a model to ask a sequence of attribute‑focused Yes/No questions, receiving Yes/No/Unsure feedback and penalizing invalid actions with Skip, thereby stressing question selection under uncertainty, constraint tracking, and fine‑grained discrimination. Using the Guess My Preferred Dress task, the study shows that final‑answer accuracy alone overstates performance, as models may guess correctly without verified evidence, waste turns, or violate the protocol, highlighting the need for combined visual discrimination, informative questioning, and robust protocol adherence.
By Min Wang, Ata Mahjoubfar
Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,...
arXiv:2609.24565v2 Announce Type: replace
Abstract: A connectome-constrained model of the fly visual system, optimized for motion and then frozen, can be driven over architectural drawings by prescri...
By Dmitry Kuklev
arXiv:2609.05762v1 Announce Type: cross
Abstract: UAV image collections contain spatially and temporally related frames, yet semantic-segmentation benchmarks commonly split them at image level. Such...
By Tarek Rahman, Nazim-E-Alam, Md Kishor Morol, Jannatun Noor
arXiv:2609.08914v2 Announce Type: replace
Abstract: Pixel-level annotation of fixed traffic-camera imagery is expensive, while crosswalk models trained from street-level imagery face a substantial vi...
By Abdirashid Omar, Jonghyuk Park
arXiv:2609.38111v1 Announce Type: new
Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when...
By Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng
LiteSearch‑VL demonstrates that distilling released agent trajectories into small vision‑language backbones can transfer the agent’s behavioral contract, enabling a 2B model to produce usable answers in 28.4% of cases on multimodal benchmarks. The approach uses parameter‑efficient LoRA adapters and synthetic step‑level preferences derived from GPT‑5 hard negatives to refine tool use and query quality. While synthetic preference learning and tool distillation provide incremental improvements, the main bottleneck identified is answer verification rather than search depth.
By Saeed Khaki, Nima Safaei, Kamal Ginotra
arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.
By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang