Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
arXiv:2606. 00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.
Tool use, function calling, orchestration and the protocols that let models act rather than only answer.
arXiv:2606. 00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.
arXiv:2606. 00860v1 Announce Type: cross Abstract: Self-report questionnaires remain the prevailing tool for probing the psychological states of persona-conditioned agents (PC-Agents).
arXiv:2512. 16167v3 Announce Type: replace-cross Abstract: Decentralized LLM-based multi-agent service economies face three vulnerabilities that undermine traditional trust mechanisms: reduced cost of fraud, difficulty in evaluating service quality, and instability of service content.
arXiv:2606. 01722v1 Announce Type: cross Abstract: For decades, distributed systems have typically assumed that correct participants execute protocol-specified behavior with stable, externally defined, and deterministic semantics.
arXiv:2606. 00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts.
arXiv:2606. 00857v1 Announce Type: cross Abstract: Accurate and reliable vehicle trajectory prediction is essential for safe autonomous driving.
arXiv:2606. 00822v1 Announce Type: cross Abstract: Skill-based LLM agents increasingly rely on long procedural documents, but full-document prompting wastes tokens and dilutes information critical to execution.
arXiv:2606. 01229v1 Announce Type: new Abstract: During green building design, computer-aided energy assessment is widely used to improve efficiency and achieve overall optimization.
arXiv:2606. 00059v1 Announce Type: cross Abstract: Informative excitation signals are critical for accurate system identification of mechatronic systems, yet classical system identification (SI) approaches require expert knowledge and hand-crafted signal design to respect hardware safety constraints, limiting their generalizability.
arXiv:2606. 00610v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become an essential method for mitigating hallucinations in Large Language Models (LLMs) by leveraging external knowledge.
arXiv:2606. 01316v1 Announce Type: new Abstract: Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces.
arXiv:2606. 01444v1 Announce Type: new Abstract: Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers are typed.
arXiv:2606. 00590v1 Announce Type: cross Abstract: Agentic search systems iteratively interact with retrieval models to answer complex queries.
arXiv:2606. 00561v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) offers a promising route to real-time power grid operation, yet large neural policies are costly to evaluate, hard to deploy on constrained hardware, and opaque to operators.
arXiv:2602. 06448v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based scientific agents have accelerated scientific discovery, yet they often suffer from significant inefficiencies due to adherence to fixed initial priors.
arXiv:2606. 01528v1 Announce Type: new Abstract: In open-ended environments, exploration is fundamental for autonomous agents, yet current language model agents struggle with this.
arXiv:2606. 01725v1 Announce Type: new Abstract: Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed outcomes.
arXiv:2605. 15229v3 Announce Type: replace-cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a described issue.
arXiv:2606. 01803v1 Announce Type: new Abstract: The explosive growth of Text-to-Image (T2I) models, from large-scale versions to lightweight, real-time ones, now faces diminishing marginal returns from single-model scaling.
arXiv:2606. 00104v1 Announce Type: cross Abstract: Foundation models are increasingly used to drive autonomous systems, yet existing approaches either keep the model in a tight control loop, raising latency and hallucination risk, or compile natural language into opaque end-to-end policies that are hard to explain, constraint and require domain-specific datasets and fine-tuning.