Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.
Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm.
Adapting CLIP for zero-shot sketch-based image retrieval (ZS-SBIR) via prompt learning faces a fundamental tension: the model must bridge the sketch-photo domain gap through task-specific adaptation, yet the added flexibility risks overfitting to seen training categories and eroding CLIP's zero-shot generalization. We present SeCo-SBIR, a semantically consistent prompt learning framework that resolves this tension from both sides.
Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain has not been systematically assessed.
arXiv:2608. 00107v1 Announce Type: new Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure.
arXiv:2608. 00850v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) have emerged as a versatile approach for solving nonlinear partial differential equations (PDEs), yet achieving high accuracy efficiently using these techniques remains challenging for high-dimensional or multiscale systems.
arXiv:2608. 00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models.
arXiv:2608. 01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count ceiling bounded by release cadence.
arXiv:2608. 01074v1 Announce Type: new Abstract: Tabular data is used extensively in many real-world use cases.
arXiv:2608. 01133v1 Announce Type: new Abstract: Evaluating Multi-Agent Reinforcement Learning (MARL) policies in autonomous driving fundamentally relies on extrinsic statistical indicators (e.
arXiv:2608. 01160v1 Announce Type: new Abstract: Topological neural networks (TNNs) enable leveraging high-order structures on graphs (e.
arXiv:2608. 01184v1 Announce Type: new Abstract: Data-free continual model merging must incorporate a stream of specialized models while retaining both pretrained general knowledge and previously acquired tasks, without access to task data.
arXiv:2608. 01303v1 Announce Type: new Abstract: Symbolic alpha factor discovery can score a completed expression, but it provides no direct label for the structural decisions that produced it.
arXiv:2608. 01616v1 Announce Type: new Abstract: Competitive analysis is central to the study of online algorithms, but upper bounds are often highly problem-specific.
arXiv:2608. 01745v1 Announce Type: new Abstract: Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation.
arXiv:2608. 01793v1 Announce Type: new Abstract: Unified anomaly detection requires modeling highly heterogeneous normal data without access to anomalous samples.
arXiv:2608. 01968v1 Announce Type: new Abstract: Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks.
arXiv:2608. 02229v1 Announce Type: new Abstract: Classical neural networks frequently produce overconfident predictions on ambiguous or out-of-distribution (OOD) data, a liability that grows with each AI system deployed in safety-critical real-world scenarios.
arXiv:2608. 02352v1 Announce Type: new Abstract: Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes.
arXiv:2608. 02305v1 Announce Type: new Abstract: Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget.