Just A Rather Very Intelligent Spoken Agent
arXiv:2607. 16610v1 Announce Type: new Abstract: Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
arXiv:2608. 14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce.
arXiv:2607. 16610v1 Announce Type: new Abstract: Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
arXiv:2606. 03103v1 Announce Type: new Abstract: Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information and users provide additional instructions, clarifications, feedback, or corrections as the task progresses.
arXiv:2607. 26300v1 Announce Type: cross Abstract: AI agents are increasingly adept at tackling complex, long-running tasks.
arXiv:2607. 23678v1 Announce Type: new Abstract: Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use.
arXiv:2509.08494v2 Announce Type: replace-cross Abstract: As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures....
arXiv:2608. 05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time.
AnyAct introduces a universal action layer that consolidates diverse tool capabilities into a self‑evolving action space for AI agents operating in open‑world environments. It tackles the scale dilemma, tool non‑stationarity, and heterogeneous feedback by using hierarchical progressive retrieval and test‑time reliability evolution, while a heterogeneous observation grounding module unifies multi‑modal feedback. Evaluations on LiveMCPBench and the newly created OSMCP benchmark show state‑of‑the‑art performance, with significant gains in task success rate and reduced execution steps, especially for models with limited native capabilities.
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
arXiv:2606. 30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the corresponding interactions on a running application, narrate them coherently, and answer questions in real time.
arXiv:2609.08977v3 Announce Type: replace-cross Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
The technical report introduces Gander, an end‑to‑end model that integrates omni perception, real‑time interaction, and agentic capabilities into a single framework. Unlike traditional turn‑based systems, Gander continuously processes streaming inputs from video, speech, and text, enabling natural full‑duplex interaction in both everyday conversations and workflow‑oriented scenarios. Its architecture features a Cerebellum‑Brain collaboration—where the Cerebellum handles real‑time interaction and omni conversational tasks while the Brain manages complex reasoning—and a streaming Thinker‑Talker design that flattens inputs and outputs into an ordered token stream for low‑latency, continuous dialogue. Evaluations across conversational ability, omni understanding, interactive capability, and agentic intelligence show that Gander matches state‑of‑the‑art open‑source models in spoken dialogue while maintaining robust performance in noisy, multi‑party, and backchannel environments.