Sherpa: Teaching LLMs to Teach Adaptively
arXiv:2610.08778v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach...
Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.
arXiv:2610.08778v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach...
arXiv:2610.05094v1 Announce Type: cross Abstract: Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design...
arXiv:2610.06892v1 Announce Type: cross Abstract: Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are l...
arXiv:2610.06927v1 Announce Type: cross Abstract: The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free r...
The paper introduces APEX, an active defense for large language model agents that protects against indirect prompt injection by enforcing safety at execution boundaries. APEX uses an evidence‑gated prevention contract and deception‑based exposure to ensure that only authorized effects, endorsed by the task, are executed. Evaluation shows APEX achieves near‑zero attack success across multiple benchmarks and capability‑unit types, outperforming 13 baseline defenses.
arXiv:2610.06977v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-sou...
arXiv:2610.06993v1 Announce Type: cross Abstract: Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However,...
arXiv:2610.06996v1 Announce Type: cross Abstract: Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memor...
arXiv:2610.07043v1 Announce Type: cross Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sam...
arXiv:2610.07062v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these response...
arXiv:2610.07098v1 Announce Type: cross Abstract: Large transformer-based models increasingly depend on multi-GPU execution, which requires frequent collective communication among GPUs. Existing comm...
arXiv:2610.07115v1 Announce Type: cross Abstract: The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging ea...
PlaySuite is a large-scale benchmark that evaluates interactive visual intelligence by using over 5,000 open-source video games from platforms like PyWeek and itch.io. The benchmark covers diverse game engines (Pygame, HTML5, Godot, Unity) and introduces a unified closed-loop interaction framework and a Video-LLM-as-a-judge protocol to standardize progress measurement. Evaluation of fourteen recent models shows a perception-action gap, with strong reasoning but poor sustained progress, spatial grounding, action execution, and self-correction.
The paper introduces TraceDSE, an agentic design space exploration framework for jointly mapping AI inference workloads to heterogeneous edge SoCs and configuring each processing unit. Unlike traditional black-box optimization, TraceDSE uses a proposer‑critic loop powered by large language models and enriched with system execution traces to identify bottlenecks and refine design choices. Experiments on an Intel Meteor Lake SoC show that TraceDSE outperforms state‑of‑the‑art evolutionary and Bayesian methods, improving Pareto frontier hypervolume by up to 68% while reducing hardware evaluations by 6–9×.
arXiv:2610.07207v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability probl...
arXiv:2610.07224v1 Announce Type: cross Abstract: Clinical notes capture most of what is documented about a patient's care, but they cannot be used for research until protected health information (PH...
The paper proposes the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result in enterprise AI agents. By gating retrieval with a policy that requires authorization for every column touched, the authors prove that sensitive columns cannot be leaked through derived results, achieving up to 90% lineage completeness to eliminate leakage. Experiments show lineage‑gated retrieval removes 18.8‑25.5% of cross‑department leakage while maintaining 81.5‑82.6% memory reuse with minimal overhead, and a real‑agent proof‑of‑concept demonstrates zero leaks over multiple interactions.
arXiv:2610.07276v1 Announce Type: cross Abstract: Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verificatio...
The paper introduces FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit continuous integration workflow. FlowAgent uses a ReAct-style generate-and-validate loop with strict latency and quality filters, and was evaluated on 195 real-world failures with a 67.18% accuracy rate. After deployment, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554, and received positive feedback from interviews.
arXiv:2610.07298v1 Announce Type: cross Abstract: Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources...