Reported Confidence in LLMs Tracks Commitment More Than Correctness
arXiv:2606. 29490v1 Announce Type: cross Abstract: Confidence is an estimate of the probability that a chosen answer is correct.
Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.
arXiv:2606. 29490v1 Announce Type: cross Abstract: Confidence is an estimate of the probability that a chosen answer is correct.
arXiv:2408. 16028v4 Announce Type: replace-cross Abstract: Supervised-learning-based vulnerability detectors often fall short due to limited labelled training data.
arXiv:2606. 29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets.
arXiv:2509. 07123v2 Announce Type: replace-cross Abstract: Generalized extreme value models capture dependence among choice alternatives in discrete choice modeling, but require this dependence to be predefined, symmetric, and shared uniformly across individuals.
arXiv:2602. 04940v2 Announce Type: replace Abstract: Deep learning has emerged as a transformative tool for the neural surrogate modeling of partial differential equations (PDEs), known as neural PDE solvers.
arXiv:2606. 30430v1 Announce Type: cross Abstract: The increasing connectivity of modern vehicles has made securing in-vehicle communication networks a critical challenge.
arXiv:2606. 29758v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training.
arXiv:2606. 29400v1 Announce Type: cross Abstract: In computer graphics, visual content is continuously warped, zoomed and resampled.
arXiv:2606. 28841v1 Announce Type: cross Abstract: Large language models are increasingly capable of mathematical reasoning, but the proofs they generate are often unreliable and hard to verify.
arXiv:2606. 29748v1 Announce Type: new Abstract: The application of graph data in numerous disciplines raises the need for gathering and analyzing huge volumes of data, some of which is private and sensitive.
arXiv:2606. 29328v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) typically treats context selection as ranking chunks against a single query embedding.
arXiv:2606. 29733v1 Announce Type: cross Abstract: Organizations that cannot send data to a cloud API increasingly ask: how good is Text-to-SQL if the model must run on-premises on open weights, and which popular accuracy "recipes" are worth their compute?
arXiv:2512. 01461v2 Announce Type: replace Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training.
arXiv:2512. 11529v3 Announce Type: replace Abstract: Recommendation system delivers substantial economic benefits by providing personalized predictions.
arXiv:2606. 29788v1 Announce Type: new Abstract: When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success.
arXiv:2606. 30358v1 Announce Type: cross Abstract: We design an algorithm for learning the coefficients of an $n$-qubit constant-local Lindbladian to $\varepsilon$ error with $O(g d^2 \log(n) / \varepsilon^2)$ total evolution time, where $g$ is the single-site energy and $d$ is the (approximate) degree of the interaction graph.
arXiv:2606. 30104v1 Announce Type: new Abstract: Electroencephalography (EEG) foundation models aim to learn generalizable representations from large-scale brain recordings.
arXiv:2606. 29280v1 Announce Type: cross Abstract: We identify intervention bias as a previously unquantified failure mode of zero-shot large-language-model (LLM) educational advisory agents: without task-specific training, they recommend action when a hindsight-optimal oracle policy mandates inaction.
arXiv:2606. 29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data.
arXiv:2606. 30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M).