Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
arXiv:2608. 09444v1 Announce Type: new Abstract: A main promise of looped language models (LMs) is depth-adaptive inference.
Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.
arXiv:2608. 09444v1 Announce Type: new Abstract: A main promise of looped language models (LMs) is depth-adaptive inference.
arXiv:2608. 09888v1 Announce Type: cross Abstract: We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning.
arXiv:2608. 01548v2 Announce Type: replace Abstract: Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--environment complex can execute, and which possibilities can be realized with finite resources.
arXiv:2606. 07998v3 Announce Type: replace-cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs), raise concerns over the interpretability, safety and sustainability of these large and opaque AI models.
arXiv:2608. 09292v1 Announce Type: new Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories.
arXiv:2608. 08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds the one-time cost of fine-tuning.
arXiv:2608. 09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks.
arXiv:2608. 09742v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.
arXiv:2603. 28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs.
arXiv:2410. 01227v2 Announce Type: replace-cross Abstract: In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of their experiences.
arXiv:2506. 02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering.
arXiv:2504. 08811v3 Announce Type: replace Abstract: Modern learning systems often struggle with joint learning across diverse scenarios and immediate adaptation to new ones, because they rely heavily on the scenario-dependent absolute data-label representations.
arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.
arXiv:2606. 22357v2 Announce Type: replace-cross Abstract: Language models are widely used in assistant settings, where controlling behavioral attributes is often essential.
arXiv:2607. 11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration.
arXiv:2605. 12530v2 Announce Type: replace-cross Abstract: LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks.
arXiv:2608. 08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models (LLMs).
arXiv:2608. 09348v1 Announce Type: new Abstract: Density estimation underlies many unsupervised tasks on tabular data such as anomaly detection, out-of-distribution detection, and data augmentation.
arXiv:2608. 08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored.
arXiv:2608. 08652v1 Announce Type: cross Abstract: We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it.