EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents
arXiv:2605. 12887v2 Announce Type: replace-cross Abstract: Web-enabled LLM agents are changing how online information influences search outcomes.
arXiv:2605. 12887v2 Announce Type: replace-cross Abstract: Web-enabled LLM agents are changing how online information influences search outcomes.
arXiv:2608.22697v1 Announce Type: new Abstract: Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and...
arXiv:2608.23045v1 Announce Type: new Abstract: Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent h...
arXiv:2606. 07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end.
The paper introduces AgentX-Model, a dual‑agent framework that links proposal development with model experimentation in industrial recommender systems. The Research Agent drafts proposals from literature and prior findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework iteratively selects starting implementations and formulates new research questions, organizing work into Reproduce, Follow‑up, Composition, and Diagnose actions. Across production evaluations, most experiments exceeded business baselines, with recent A/B tests showing significant gains in acquisition efficiency, advertising spend, and watch time while reducing computational cost.
The paper discusses how enterprises increasingly deploy AI coding agent harnesses, often purchased from vendors like Anthropic or OpenAI, and how these harnesses dictate model choice, prompt handling, and cost. It introduces a fast, customizable routing system that classifies prompts and strategically routes them to minimize expensive model usage, achieving 14–21% cost savings in a simulated 10,000-seat enterprise. The study also evaluates risks across twenty harnesses, highlights vendor dependence, and proposes an internal control plane for future harness ownership decisions.
The "Era by Eon" benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from generated company data. While top models can answer most questions, the benchmark introduces eight new templates that rely on hidden facts not explicitly stated in any document, making the task harder. Evaluation of 12 agents shows that only the best agent correctly answers 18 of 24 attempts, with many questions remaining largely unsolved.
arXiv:2607.10198v2 Announce Type: replace Abstract: Search APIs expose ranked snippets, URLs, and metadata on which agents decide whether to answer, search again, or fetch pages. We evaluate these in...
arXiv:2607. 10286v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value.
The Era by Eon benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from a company’s data. In the original benchmark, top models answered 22–25 of 27 questions, barely distinguishing performance. The updated benchmark adds eight templates that rely on hidden facts not explicitly stated in any question or document, forcing agents to infer information from indirect data. Twelve agents were evaluated, with the best achieving 18 of 24 correct answers, while the hardest questions—requiring selection among similar records—were answered correctly only 1 out of 84 attempts across all agents.
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
arXiv:2606. 30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified.