arXiv Computation and Language
Sep 11

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

The paper announces VectraYX‑Vision‑1B, a Spanish/LATAM cybersecurity vision‑language model that uses a natively‑trained visual tower (Qwen2‑VL) and a frozen decoder. The authors report that transplanting the visual tower improves the failing 8‑nibble address field from 0.00 to 0.81 exact, using a coarser token budget than 2×2 tiling, and that a second pre‑registered field shows 63% contamination and is demoted. The model’s B6/B7 tool‑id remains at floor, and code/checkpoints are available on Hugging Face.

By Juan S. Santillana
arXiv Computer Vision
Sep 24

Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection

The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.

By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim