Beyond Component Testing: Validating Agentic AI Systems
arXiv:2607. 29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation.
Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.
arXiv:2607. 29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation.
arXiv:2607. 29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases.
arXiv:2607. 29456v1 Announce Type: cross Abstract: Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference.
arXiv:2607. 29283v1 Announce Type: cross Abstract: Training large language models (LLMs) to write register-transfer level (RTL) requires large corpora of paired specifications and code, and such data is scarce enough that most public corpora are now synthesized.
arXiv:2602. 02685v3 Announce Type: replace Abstract: Decentralized Diffusion Models (DDMs) route denoising through experts trained independently on disjoint data clusters, which can strongly disagree in their predictions.
arXiv:2607. 28946v1 Announce Type: new Abstract: Despite its many benefits, widespread access to individuals' personal data also causes severe privacy concerns for consumers, companies, and policymakers.
arXiv:2607. 29053v1 Announce Type: new Abstract: Standard model comparison is global, aggregating losses across the covariate space to declare a single winner.
arXiv:2510. 24598v2 Announce Type: replace Abstract: Current quantum machine learning approaches often face challenges balancing predictive accuracy, robustness, and interpretability.
arXiv:2605. 07663v2 Announce Type: replace-cross Abstract: Data valuation methods allocate payments and audit training data's contribution to machine-learning pipelines; however, they often assume passive contributors.
arXiv:2605. 04215v3 Announce Type: replace-cross Abstract: Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead to significant throughput advantages and superior GPU utilization over the traditional autoregressive paradigm.
arXiv:2607. 28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning.
arXiv:2603. 13683v4 Announce Type: replace-cross Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts.
arXiv:2604. 26051v2 Announce Type: replace-cross Abstract: The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promising approach for operational flood monitoring.
arXiv:2605. 07699v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations.
arXiv:2510. 17136v2 Announce Type: replace Abstract: The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models.
arXiv:2607. 29240v1 Announce Type: cross Abstract: In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state.
arXiv:2607. 29586v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid.
arXiv:2607. 29624v1 Announce Type: cross Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback.
arXiv:2607. 28903v1 Announce Type: cross Abstract: Variance-based global sensitivity analysis (GSA) plays a key role in uncertainty quantification by identifying the contributions of uncertain inputs to the variability of the model response.
arXiv:2510. 06945v2 Announce Type: replace Abstract: Motivated by the growing interest in quantum machine learning, in particular quantum neural networks (QNNs), we study how recently introduced evaluation metrics based on the Fisher information matrix (FIM) are effective for predicting their training and prediction performance.