We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1. 04B Spanish/LATAM security decoder via an MLP.
The paper announces VectraYX‑Vision‑1B, a Spanish/LATAM cybersecurity vision‑language model that uses a natively‑trained visual tower (Qwen2‑VL) and a frozen decoder. The authors report that transplanting the visual tower improves the failing 8‑nibble address field from 0.00 to 0.81 exact, using a coarser token budget than 2×2 tiling, and that a second pre‑registered field shows 63% contamination and is demoted. The model’s B6/B7 tool‑id remains at floor, and code/checkpoints are available on Hugging Face.
By Juan S. Santillana
The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.
By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,...
arXiv:2606. 14629v1 Announce Type: cross Abstract: Verifier-driven self-DPO is a common recipe for self-improving production visual-language models.
By Jianzhe Lin
The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.
By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim