Measuring Open-Source Llama Nemotron Models on DeepResearch Bench
Related stories
Welcome the NVIDIA Llama Nemotron Nano VLM to Hugging Face Hub
Llama 2 on Amazon SageMaker a Benchmark
BenchMIRT: What are LLM benchmarks actually measuring?
Machine Learning and Deep Learning for Exoplanet Detection and Atmospheric Characterization with JWST and the Upcoming Ariel Mission
arXiv:2606. 23766v1 Announce Type: cross Abstract: The detection and atmospheric characterization of exoplanets have entered a new data-intensive era driven by the James Webb Space Telescope and the upcoming Ariel mission.
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports
Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.
Open-source DeepResearch – Freeing our search agents
StackLLaMA: A hands-on guide to train LLaMA with RLHF
From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates
Understanding How DeepSeek's Flagship Open-Weight Models Evolved
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs
OlmoEarth v1.2: A more efficient family of OlmoEarth models
arXiv:2605. 20804v2 Announce Type: replace-cross Abstract: We present a set of improvements to the OlmoEarth family.
Announcing our partnership with the Republic of Korea
Google DeepMind and Korea partner to accelerate scientific breakthroughs using frontier AI models

