Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study
arXiv:2608. 13568v1 Announce Type: cross Abstract: Coding agents spend most of their context budget on retrieval.
arXiv:2608. 13568v1 Announce Type: cross Abstract: Coding agents spend most of their context budget on retrieval.
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.
Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the technical-support team of AFIR, the Romanian Agency for Financing Rural Investments, built and operated under three hard constraints: zero data egress (no external API calls, even for synthetic data), a read-only mandate (the assistant drafts, humans execute), and a single 8 GB consumer laptop as the only development and training machine.
The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.
MemToC is a controlled benchmark that tests how large language models resolve conflicts between their internal memory and tool outputs. It contains 6,504 episodes built from 542 factual questions, each paired with a model‑generated closed‑book answer and a tool return whose correctness is known, creating four distinct source‑correctness scenarios. Across five 7‑9B open‑weight models, tool responses overwhelmingly dominate closed‑book answers, and only a minority of instruction‑tuned models correctly retain a verified answer when the tool is wrong, while most follow a correct tool or repeat a wrong tool.
The paper investigates why large vision‑language models sometimes misclassify harmful memes, attributing failures to either missing internal evidence or poor routing of evidence to the output. Using sparse autoencoders, role‑conditioned probes, and causal interventions on Gemma‑3 and Qwen3.5, the authors show that sparse readouts consistently outperform native predictions across six harmful content benchmarks, revealing a readout gap that is largely due to routing rather than representation. The study also demonstrates that calibration‑only routing recovers most of the performance gap and that the issue persists across languages and is not solely driven by OCR signals.
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
CoHyDE is an iterative co‑training framework that jointly trains a dense encoder and an LLM rewriter for tool retrieval from large API catalogs. The encoder is fine‑tuned with InfoNCE on catalog‑style hypothetical descriptions generated by the rewriter, while the rewriter is preference‑aligned via DPO against the encoder’s retrieval scores. On a 10k‑tool subset of ToolBench, three rounds of CoHyDE outperform the best single‑component baseline by 2.5 pp NDCG@5 on standard queries and 6.3 pp on vague queries, with the largest gains on the hardest vague tier.
The paper investigates why large‑language‑model coding agents rarely request a second chunk of tool output, focusing on the precision‑at‑1 rate ($p_1$) of the gold item appearing first in the first chunk. In a benchmark of 500 software‑engineering tasks, the authors compare six value functions and find that increasing $p_1$ does not systematically improve downstream accuracy; the agent can recover the correct answer from any position within the chunk. Adding file‑metadata signals to a keyword scorer actually reduces $p_1$, while a parameter‑free keyword scorer improves $p_1$ but still fails to boost overall accuracy.
Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.
arXiv:2609.14144v1 Announce Type: cross Abstract: A transformer language model is trained to respond to any prompt, but each deployment asks only a narrow range of questions: a support assistant sees...
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.