The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
arXiv:2608. 11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem.
Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.
arXiv:2608. 11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem.
arXiv:2608. 11623v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting.
arXiv:2608. 11216v1 Announce Type: new Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments.
arXiv:2608. 12262v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration.
arXiv:2608. 12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior.
arXiv:2608. 11801v1 Announce Type: new Abstract: Multivariate time-series anomaly prediction aims to identify whether and when anomalies will occur over a future horizon from historical observations.
arXiv:2608. 11492v1 Announce Type: cross Abstract: IoT firmware vulnerability detection remains challenging due to heterogeneous firmware ecosystems, resource-constrained platforms, and limitations in existing benchmarks.
arXiv:2608. 11286v1 Announce Type: cross Abstract: Cyberattack detection in electric vehicle charging infrastructure is complicated by legitimate post-activation revisions to requested energy and departure time.
arXiv:2608. 11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity.
arXiv:2608. 11283v1 Announce Type: cross Abstract: Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity.
arXiv:2608. 11675v1 Announce Type: new Abstract: Coupon campaigns seek to lift both conversion and revenue, but gross merchandise value (GMV) follows a deterministic funnel from conversion to conditional order value and is zero-inflated and heavy-tailed.
arXiv:2608. 11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon.
arXiv:2608. 11238v1 Announce Type: new Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests.
arXiv:2608. 11616v1 Announce Type: new Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation.
arXiv:2608. 11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al.
arXiv:2608. 12220v1 Announce Type: cross Abstract: Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning.
arXiv:2608. 11225v1 Announce Type: new Abstract: AI "personality clones" force a re-examination of personal identity in operational terms.
arXiv:2608. 11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs).
arXiv:2608. 11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.
arXiv:2608. 11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs.