arXiv Machine Learning

When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

arXiv:2605. 09030v2 Announce Type: replace-cross Abstract: Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity score for text-to-image and style-imitation evaluation.

arXiv AI
Jun 16

StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer

arXiv:2605. 00924v2 Announce Type: replace-cross Abstract: AI-generated content (AIGC) detectors are increasingly deployed in high-stakes settings such as academic integrity screening, yet their reliability rests on a fundamental paradox: as language models are trained on human-written corpora, the statistical boundary between AI and human writing will inevitably dissolve as models improve.

By Guantian Zheng
arXiv Machine Learning
Aug 26

The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

The paper investigates the blending ratio used in few‑shot adaptation of vision‑language models, which combines a zero‑shot text prototype with the mean of labeled image features. It shows that the theoretically optimal ratio—derived from a closed‑form mean‑squared error minimizer—does not align with the ratio that actually maximizes performance, falling short by an average of 8.5 points. Moreover, a leave‑one‑out estimate on the support set achieves near‑oracle performance, and validation‑free linear probes outperform even oracle‑tuned blends, indicating that the hyperparameter can be set near‑optimally without external validation data.

By Liangzhi Li, Bowen Wang, Yiming Qian, Thorsten Neumann, Xia Xie, Guangshun Li
arXiv Machine Learning
Sep 24

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.

By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
arXiv AI
2d ago

On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence

The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.

By Vinay Kumar Chaganti
arXiv Computer Vision
Sep 2

Can Scene Text Recognition Read Rare Compositions?

The paper reports that scene text recognition models, while achieving 89–97% accuracy on standard benchmarks, perform significantly worse on rare word–trigram combinations, with a 10–18 point drop in accuracy at the rare‑word/rare‑trigram corner across multiple languages and models. Scaling the vision backbone improves overall accuracy but does not alleviate this corner‑specific deficit. The authors identify the autoregressive decoder’s lexical prior as the root cause and show that architectural changes—specifically moving from autoregressive to CTC decoding—yield the largest improvement for these rare compositions.

By Genpei Zhang
arXiv AI
4d ago

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.

By Aditya Sharma, Divya Saxena
arXiv Machine Learning
Sep 11

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.

By Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige
arXiv Machine Learning
Sep 14

PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition

The paper introduces PA-CDM, a new position‑aware character detection matching metric for handwritten mathematical expression recognition that improves on existing render‑based and tree‑edit metrics by incorporating position‑forest encoding and divergence‑level weighting. It also presents StructPerturb v2.0, a benchmark of 1,340 controlled perturbation pairs, and a cross‑metric consistency protocol that includes a sensitivity matrix, a human study, and LLM‑judge calibration. In a six‑annotator study, PA‑CDM achieves the highest correlation with human judgments (Spearman rho = 0.9535) among seven automatic metrics, approaching the performance of a costly LLM judge while remaining deterministic and cost‑free.

By Shiliang Luo (East China Normal University)