We present Seed1.8, a foundation model aimed at generalized real-world agency: going beyond single-turn prediction to multi-turn interaction, tool use, and multi-step execution. Seed1.8 keeps strong LLM and vision-language performance while supporting a unified agentic interface-search, code generation and execution, and GUI interaction. For deployment, it offers latency- and cost-aware inference, including configurable thinking modes and optimized visual encoding for images and video. We report evaluations on standard benchmarks and application-aligned workflows spanning foundational skills, multimodal understanding, and agentic behavior. Seed1.8 is released to support further research and development on interactive, real-world use cases.
核心贡献 · Key contributions
提出 Seed1.8,一个面向通用现实世界智能体的基础模型,支持多轮交互、工具使用和多步执行。 Proposes Seed1.8, a foundation model for generalized real-world agency with multi-turn interaction, tool use, and multi-step execution.
将强大的 LLM 和 VLM 能力与统一的智能体接口集成,支持搜索、代码生成和 GUI 交互。 Integrates strong LLM and VLM capabilities with a unified agentic interface for search, code generation, and GUI interaction.
引入可配置的思考模式(无思考、低/中/高思考),实现延迟和成本感知的推理。 Introduces configurable thinking modes (no_think, think-low/medium/high) for latency- and cost-aware inference.
在多个智能体基准测试(包括 GAIA、OSWorld 和 BrowseComp)上达到最先进性能。 Achieves state-of-the-art performance on multiple agentic benchmarks including GAIA, OSWorld, and BrowseComp.
在多模态任务中展现出强大的词元效率,尤其是在最小词元预算下的长视频理解方面。 Demonstrates strong token efficiency in multimodal tasks, especially long-video understanding with minimal token budgets.
通过面向现实的基准测试,在基础、多模态和智能体能力方面提供全面评估。 Provides comprehensive evaluations across foundational, multimodal, and agentic capabilities with real-world-oriented benchmarks.
局限 · Limitations
在学科视频知识基准(VideoMMMU、MMVU)上的表现落后于 Gemini-2.5/3-Pro。 Performance on disciplinary video knowledge benchmarks (VideoMMMU, MMVU) lags behind Gemini-2.5/3-Pro.
在 TOMATO 基准上的运动感知与人类表现仍有显著差距。 Motion perception on TOMATO benchmark still has a significant gap compared to human performance.
在 SWE-bench Verified 等智能体编码基准上表现持平但未超越顶级模型。 Agentic coding benchmarks like SWE-bench Verified show performance on par but not surpassing top models.
面向高价值应用的内部基准可能无法完全泛化到所有现实场景。 Internal benchmarks for high-value applications may not fully generalize to all real-world scenarios.
安全评估依赖内部基准和开源测试,未涵盖所有风险类型。 Safety evaluations rely on internal benchmarks and open-source tests; coverage of all risk types is not exhaustive.