Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
We present Seed2. 0, a model series that takes a meaningful step toward solving complex, real-world tasks.
arXiv:2607. 00248v1 Announce Type: new Abstract: We present Seed2.
We present Seed2. 0, a model series that takes a meaningful step toward solving complex, real-world tasks.
arXiv:2606. 14397v1 Announce Type: new Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities.
InternBootcamp is an open‑source framework that offers over 1,000 domain‑diverse task environments for large language model (LLM) reasoning research. It introduces Bootcamp‑Eval, an automatically generated benchmark for comprehensive performance assessment. Experiments show that training on InternBootcamp significantly improves reasoning performance, with a 32B model achieving state‑of‑the‑art results on Bootcamp‑Eval and other established benchmarks, demonstrating that scaling the number of training tasks yields consistent gains.
arXiv:2606. 29929v1 Announce Type: new Abstract: Distilling historical trajectories into reusable experience to enhance future problem-solving has become a focal point of recent LLM research.
arXiv:2508.16821v2 Announce Type: replace Abstract: We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinfo...
arXiv:2609.14473v1 Announce Type: new Abstract: Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical...
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
arXiv:2606.14397v4 Announce Type: replace Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capab...
arXiv:2602. 10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors.
arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.
arXiv:2608.28638v1 Announce Type: new Abstract: Agent skills are portable packages of instructions and resources an agent consults at deployment. Self-evolving them fails in two ways today. First, sk...
arXiv:2512. 11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space.