MELLON - Multimodal Enhanced LLM for Online Navigation
arXiv:2608. 09121v1 Announce Type: new Abstract: Web navigation agents are capable of addressing various types of tasks on different websites.
Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.
arXiv:2608. 09121v1 Announce Type: new Abstract: Web navigation agents are capable of addressing various types of tasks on different websites.
arXiv:2602. 00549v2 Announce Type: replace Abstract: While Monte Carlo Tree Search (MCTS) shows promise in Large Language Model (LLM) based Automatic Heuristic Design (AHD), it suffers from a critical over-exploitation tendency under the limited computational budgets required for heuristic evaluation.
arXiv:2608. 07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction.
arXiv:2608. 07540v1 Announce Type: new Abstract: AI systems increasingly operate between flexible input representations and formal objects used by downstream tools.
arXiv:2608. 07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting).
arXiv:2608. 07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness.
arXiv:2608. 07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address.
arXiv:2608. 08189v1 Announce Type: new Abstract: LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive.
arXiv:2608. 07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task.
arXiv:2608. 08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy.
arXiv:2608. 08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks.
arXiv:2608. 08176v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation.
arXiv:2608. 08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions.
arXiv:2608. 08210v1 Announce Type: new Abstract: Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an \textbf{illusion of alignment (IoA)}.
arXiv:2608. 08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge.
arXiv:2608. 08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools.
arXiv:2608. 08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control.
arXiv:2608. 08303v1 Announce Type: new Abstract: Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks.
arXiv:2608. 07631v1 Announce Type: cross Abstract: LLM-based full-duplex voice services allow users to speak while the assistant is responding.
arXiv:2608. 08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.