Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
arXiv:2608. 02786v1 Announce Type: new Abstract: AI systems can fail silently.
arXiv:2606. 13172v2 Announce Type: replace Abstract: Learned representations are central to modern machine learning and are commonly evaluated through predictive performance, robustness, uncertainty estimation, and generalization.
arXiv:2608. 02786v1 Announce Type: new Abstract: AI systems can fail silently.
arXiv:2410. 22526v2 Announce Type: replace Abstract: To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level hazards.
arXiv:2607. 19386v1 Announce Type: new Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores.
arXiv:2607. 22797v1 Announce Type: cross Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon.
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
arXiv:2605. 06890v3 Announce Type: replace Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagnose and control.
arXiv:2605. 23055v2 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.
arXiv:2605. 27618v2 Announce Type: replace Abstract: Despite the wide use of explainability techniques to attempt to understand the behavior of Artificial Intelligence (AI), the generated explanations may not always be reliable.
arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.
arXiv:2607. 28567v1 Announce Type: cross Abstract: Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times.
arXiv:2606. 07303v1 Announce Type: new Abstract: Representation learning is central to modern machine learning, enabling transitions from handcrafted features to learned embeddings, latent spaces, foundation models, world models, and digital twins.
arXiv:2603. 25450v2 Announce Type: replace Abstract: Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment.