VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
arXiv:2608. 10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants.
Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.
arXiv:2608. 10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants.
arXiv:2608. 11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways.
arXiv:2509. 19696v4 Announce Type: replace-cross Abstract: Learning-based methods excel at robot motion generation but remain limited in contact-rich physical interaction.
arXiv:2602. 13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences.
arXiv:2604. 13201v2 Announce Type: replace-cross Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging.
arXiv:2608. 10584v1 Announce Type: new Abstract: Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery.
arXiv:2608. 10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains.
arXiv:2608. 10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them.
arXiv:2608. 10823v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging overhead.
arXiv:2608. 10729v1 Announce Type: cross Abstract: Foundation models can improve their outputs through a self-refinement process driven by external feedback.
arXiv:2608. 00422v2 Announce Type: replace Abstract: Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncertainty estimation important for abstention, human review, and adaptive compute allocation.
arXiv:2608. 11197v1 Announce Type: new Abstract: Shani et al.
arXiv:2608. 10567v1 Announce Type: new Abstract: Analytic dashboards combine coordinated views and interactions for data exploration and decision-making.
arXiv:2604. 07822v2 Announce Type: replace-cross Abstract: We study implicit reasoning, i.
arXiv:2608. 10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit.
arXiv:2608. 10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems.
arXiv:2608. 10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models.
arXiv:2608. 10375v1 Announce Type: cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions.
arXiv:2605. 17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages.
arXiv:2608. 09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting.