Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
arXiv:2608. 04265v1 Announce Type: cross Abstract: Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed.
Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.
arXiv:2608. 04265v1 Announce Type: cross Abstract: Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed.
arXiv:2608. 04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied.
arXiv:2608. 04351v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck.
arXiv:2608. 04405v1 Announce Type: cross Abstract: Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches.
arXiv:2608. 04504v1 Announce Type: cross Abstract: Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables.
arXiv:2608. 04515v1 Announce Type: cross Abstract: Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices.
arXiv:2608. 04588v1 Announce Type: cross Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents.
arXiv:2608. 04670v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions.
arXiv:2608. 04698v1 Announce Type: cross Abstract: We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expression when it exists (positive sample) and to refuse output when it does not (negative sample).
arXiv:2511. 21692v3 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation.
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
arXiv:2608. 04872v1 Announce Type: cross Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt.
arXiv:2608. 04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code.
arXiv:2608. 05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images.
arXiv:2608. 05063v1 Announce Type: cross Abstract: The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.
arXiv:2602. 11745v2 Announce Type: replace Abstract: Graph models are fundamental to data analysis in domains rich with complex relationships.
arXiv:2607. 21597v2 Announce Type: replace Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal.
arXiv:2507. 06506v2 Announce Type: replace-cross Abstract: Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine translation systems.
arXiv:2508. 18066v2 Announce Type: replace-cross Abstract: Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge.
arXiv:2602. 11506v4 Announce Type: replace-cross Abstract: The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characterization on resource-constrained edge hardware.