TriQua: Reconciling Granularity and Context in Factuality Evaluation
arXiv:2608. 05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.
Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.
arXiv:2608. 05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.
arXiv:2605. 20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs).
arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
arXiv:2608. 05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time.
arXiv:2608. 05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set.
arXiv:2608. 05157v1 Announce Type: cross Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias.
arXiv:2608. 05705v1 Announce Type: cross Abstract: Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings.
arXiv:2608. 05172v1 Announce Type: cross Abstract: The task-based framework in economics models occupations as bundles of tasks.
arXiv:2608. 05604v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time.
arXiv:2608. 05668v1 Announce Type: cross Abstract: With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial trading.
arXiv:2608. 02616v2 Announce Type: replace-cross Abstract: We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.
arXiv:2608. 05266v1 Announce Type: new Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines.
arXiv:2608. 00641v2 Announce Type: replace Abstract: Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and optimization stages.
arXiv:2605. 04970v3 Announce Type: replace-cross Abstract: Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly.
arXiv:2608. 06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibration, and Multimodal Instruction Tuning.
arXiv:2608. 05375v1 Announce Type: new Abstract: Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity.
arXiv:2608. 05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution.
arXiv:2608. 06085v1 Announce Type: new Abstract: Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random.
arXiv:2608. 05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization.
arXiv:2601. 21800v4 Announce Type: replace Abstract: We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks.