When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
arXiv:2601. 07965v2 Announce Type: replace Abstract: When a model knows when it does not know, many possibilities emerge.
Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.
arXiv:2601. 07965v2 Announce Type: replace Abstract: When a model knows when it does not know, many possibilities emerge.
arXiv:2507. 07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously.
arXiv:2606. 29613v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have recently been extended with role-based mechanisms for interpretability.
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
arXiv:2606. 29240v1 Announce Type: new Abstract: Heterogeneous graph neural networks (HGNNs) have achieved strong performance in modeling complex graph-structured data with multiple node and relation types.
arXiv:2606. 29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
arXiv:2606. 29894v1 Announce Type: cross Abstract: As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
arXiv:2606. 30190v1 Announce Type: cross Abstract: Existing domain-incremental learning (DIL) strategies call for massive amounts of data to adapt to new domains and suffer from the overfitting problem in the case of data scarcity.
arXiv:2606. 29403v1 Announce Type: cross Abstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups.
arXiv:2606. 29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search.
arXiv:2505. 12526v2 Announce Type: replace Abstract: Temporal graph networks suffer from irregular supervision in realworld dynamic graphs, as most minibatches contain few labeled events.
arXiv:2605. 16401v2 Announce Type: replace-cross Abstract: While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environmental impact, and a "one-size-fits-all" approach that ignores varying sample complexity.
arXiv:2606. 28326v1 Announce Type: cross Abstract: This research aims to solve the challenge of video retrieval from massive datasets, caused by ambiguous user queries.
arXiv:2606. 29929v1 Announce Type: new Abstract: Distilling historical trajectories into reusable experience to enhance future problem-solving has become a focal point of recent LLM research.
arXiv:2603. 25144v2 Announce Type: replace-cross Abstract: Dataset distillation (DD) compresses a large training set into a small synthetic set, reducing storage and training cost, and has shown strong results on general benchmarks.
arXiv:2510. 18989v2 Announce Type: replace Abstract: Neural operators are commonly utilized as fast surrogates for numerical solvers in PDE problems, mapping input functions to solution functions.
arXiv:2606. 28433v1 Announce Type: new Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy for learning in deployment settings.
arXiv:2601. 01569v4 Announce Type: replace Abstract: LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that struggle with long-horizon tasks due to fragile multi-turn dependencies and context drift.
arXiv:2605. 03283v2 Announce Type: replace-cross Abstract: We provide a unified theoretical analysis of Linear Discriminant Analysis with simultaneous multilabel scatter matrix formulations and Stiefel orthogonality constraints.
arXiv:2606. 28401v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that is not grounded in the image.