AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
arXiv:2607. 06624v1 Announce Type: new Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents.
Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.
arXiv:2607. 06624v1 Announce Type: new Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents.
arXiv:2607. 06948v1 Announce Type: cross Abstract: The accuracy of existing leaf-wood segmentation methods for tree point clouds varies across forest types and sites.
arXiv:2607. 06776v1 Announce Type: new Abstract: We introduce an efficient Bayesian deep ensemble method for predictive regression designed to enhance interpretability while maintaining competitive predictive performance and computational efficiency.
arXiv:2607. 06839v1 Announce Type: new Abstract: Existing NAS benchmarks (e.
arXiv:2512. 02076v2 Announce Type: replace-cross Abstract: We propose FDRMFL, a task-driven multimodal feature extraction framework for federated regression under non-IID data distributions.
arXiv:2506. 04985v2 Announce Type: replace Abstract: Large language models (LLMs) require substantial compute, and thus energy, at inference time.
arXiv:2503. 15581v2 Announce Type: replace Abstract: Real-time safety assessment is critical for ensuring the reliable operation of complex dynamic systems.
arXiv:2605. 09778v2 Announce Type: replace Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token.
arXiv:2607. 06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred.
arXiv:2509. 21160v2 Announce Type: replace-cross Abstract: With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes.
arXiv:2508. 17298v3 Announce Type: replace-cross Abstract: Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground intermediate concepts, and perform multi-step logical inference.
arXiv:2607. 07708v1 Announce Type: cross Abstract: Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization.
arXiv:2607. 06766v1 Announce Type: cross Abstract: At Amazon Prime Video, we face the critical operational challenge of managing code deployments during live events and rapid feature releases without causing service outages.
arXiv:2607. 07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not.
arXiv:2511. 13726v2 Announce Type: replace-cross Abstract: We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models.
arXiv:2607. 07077v1 Announce Type: cross Abstract: Functional brain networks exhibit a hierarchical organization across ROI, community, and whole-brain levels, supporting local processing, inter-community coordination, and global integration.
arXiv:2601. 16529v4 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines.
arXiv:2607. 07318v1 Announce Type: cross Abstract: Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections.
arXiv:2607. 07178v1 Announce Type: cross Abstract: Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks.
arXiv:2605. 11325v3 Announce Type: replace-cross Abstract: Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy.