We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1. 04B Spanish/LATAM security decoder via an MLP.
The paper announces VectraYX‑Vision‑1B, a Spanish/LATAM cybersecurity vision‑language model that uses a natively‑trained visual tower (Qwen2‑VL) and a frozen decoder. The authors report that transplanting the visual tower improves the failing 8‑nibble address field from 0.00 to 0.81 exact, using a coarser token budget than 2×2 tiling, and that a second pre‑registered field shows 63% contamination and is demoted. The model’s B6/B7 tool‑id remains at floor, and code/checkpoints are available on Hugging Face.
By Juan S. Santillana
The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.
By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,...
arXiv:2606. 14629v1 Announce Type: cross Abstract: Verifier-driven self-DPO is a common recipe for self-improving production visual-language models.
By Jianzhe Lin
The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.
By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
arXiv:2609.31234v1 Announce Type: new
Abstract: Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be...
By Zhongyu Pang
The paper investigates how providing execution traces to multimodal judges in agentic video‑generation systems can bias their verdicts. On a benchmark of 109 two‑event clips, traces that falsely report successful tool calls cause large‑language‑model judges to incorrectly accept 78–90 % of failures, while contradictory traces lead to 100 % rejection of correct clips. The effect persists even when judges are instructed to consider only the video frames, indicating that the vulnerability stems from the judges’ learned trust in tool logs rather than the visual content itself.
By Jian Xu
The paper explores soft prompting for few‑shot object detection with vision‑language models, showing that optimizing a small number of continuous prompt tokens—especially when placed at the cross‑modal boundary and initialized from an empty space token—can match LoRA performance while training far fewer parameters. Soft prompting also avoids catastrophic forgetting, transfers to newer models, and can be verbalized into readable prompts. The study extends these findings to manipulation tasks, indicating that VLMs already contain much of the necessary knowledge for specialized domains, and the main challenge is learning how to ask for it.
By Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan
arXiv:2609.18084v1 Announce Type: cross
Abstract: Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to e...
By Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev, Jeffrey Ichnowski
arXiv:2609.16145v1 Announce Type: new
Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? W...
By Gautam Kishore
The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong