DeepTutor: Towards Agentic Personalized Tutoring
arXiv:2604. 26962v3 Announce Type: replace-cross Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs).
DeepEdu‑v1 is an AI‑tutoring system tailored for Vietnamese education that addresses data‑sovereignty and local curriculum alignment issues. It uses a long‑context inference engine to reduce retrieval calls and prefill latency by about 35%, and a self‑improving agentic layer that curates verified local knowledge without fine‑tuning. In deployment, DeepEdu achieves nearly twice the speed of standard vLLM serving and raises agentic accuracy from 70.0% to 79.5% on complex tasks, especially in financial reasoning and interactive‑agent benchmarks.
arXiv:2604. 26962v3 Announce Type: replace-cross Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs).
arXiv:2608. 08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge.
Combee is a new framework that scales prompt learning for self‑improving language model agents by enabling many agents to run in parallel while learning from their combined traces. It uses parallel scans, an augmented shuffle mechanism, and a dynamic batch size controller to maintain quality and reduce delay. Experiments on AppWorld, Terminal‑Bench, Formula, and FiNER show up to 17× speedup over prior methods with comparable or better accuracy at similar cost.
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.
Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the technical-support team of AFIR, the Romanian Agency for Financing Rural Investments, built and operated under three hard constraints: zero data egress (no external API calls, even for synthetic data), a read-only mandate (the assistant drafts, humans execute), and a single 8 GB consumer laptop as the only development and training machine.
arXiv:2607.26621v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendati...
arXiv:2607. 19358v1 Announce Type: new Abstract: Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm.
arXiv:2606. 03841v1 Announce Type: new Abstract: Recent progress in Large Language Model (LLM) agents has enabled promising advances in automated data science.
arXiv:2607. 20064v1 Announce Type: new Abstract: Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents.
arXiv:2608. 03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).
arXiv:2606. 17803v1 Announce Type: new Abstract: Large language models achieve strong reasoning performance by scaling inference-time compute, yet remain fundamentally stateless, discarding the rich, self-produced reasoning traces generated during this process.
arXiv:2605.05726v2 Announce Type: replace Abstract: As LLM agents are increasingly deployed with large libraries of reusable skills, selecting the right skill for a user request has become a critical...