Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
arXiv:2607. 00248v1 Announce Type: new Abstract: We present Seed2.
We present Seed2. 0, a model series that takes a meaningful step toward solving complex, real-world tasks.
arXiv:2607. 00248v1 Announce Type: new Abstract: We present Seed2.
InternBootcamp is an open‑source framework that offers over 1,000 domain‑diverse task environments for large language model (LLM) reasoning research. It introduces Bootcamp‑Eval, an automatically generated benchmark for comprehensive performance assessment. Experiments show that training on InternBootcamp significantly improves reasoning performance, with a 32B model achieving state‑of‑the‑art results on Bootcamp‑Eval and other established benchmarks, demonstrating that scaling the number of training tasks yields consistent gains.
arXiv:2606. 14397v1 Announce Type: new Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities.
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
arXiv:2609.14473v1 Announce Type: new Abstract: Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical...
arXiv:2602. 10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors.
arXiv:2508.16821v2 Announce Type: replace Abstract: We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinfo...
xDailyBench is a new benchmark comprising 248 tasks across 51 real‑life scenarios, designed to evaluate large language models on everyday professional consultation. The tasks are based on actual user requests and assessed with detailed binary rubrics that capture both explicit instructions and implicit needs inferred from context. In tests of 11 leading models, the best achieved a 75.6% task‑level score, yet all models struggled more with implicit requirements, showing gaps of at least 9 percentage points.
arXiv:2504.04711v2 Announce Type: replace Abstract: Current large language models (LLMs) are constrained by human-derived training data and limited by a single level of abstraction that impedes defin...
arXiv:2606.14397v4 Announce Type: replace Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capab...
arXiv:2606. 29929v1 Announce Type: new Abstract: Distilling historical trajectories into reusable experience to enhance future problem-solving has become a focal point of recent LLM research.
arXiv:2606. 06462v1 Announce Type: new Abstract: Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance.