arXiv AI By Tanmay Asthana, Aman Saksena, Divyansh Sahu

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

Read the original on arXiv AI →

arXiv:2605. 17554v2 Announce Type: replace Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Aug 7

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

arXiv:2608. 06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks.

By Bo Deng (Beihang University, Qwen DianJin Team, Alibaba Cloud Computing), Kang Zhou (Qwen DianJin Team, Alibaba Cloud Computing), Lifan Guo (Qwen DianJin Team, Alibaba Cloud Computing), Chongyang Tao (Beihang University), Xuanren Chen (Beihang University), Chenggang Xie (Beihang University), Renzhao Liang (Beihang University), Feng Chen (Qwen DianJin Team, Alibaba Cloud Computing), Chi Zhang (Qwen DianJin Team, Alibaba Cloud Computing)