Gemma 4 技术报告

Gemma 4 Technical Report

Thomas Mesnard Thomas Mesnard · · 2026-07-02 · arXiv:2607.02770 ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 Gemma 4,这是 Gemma 模型家族中新一代开源权重、原生多模态语言模型。为了提升计算效率和推理能力,Gemma 4 模型系列采用了密集和混合专家架构,参数量从 2.3B 到 31B 不等。除了为所有模型尺寸改进了视觉和音频编码器之外,我们还为 12B 模型提出了一种统一的、无编码器的架构,该架构直接处理原始音频和图像块。此外,我们集成了一种思考模式,使 Gemma 模型能够在响应之前生成推理轨迹。我们通过关键设计选择提高了推理速度、内存和计算效率以及长上下文能力。Gemma 4 在 STEM、多模态和长上下文基准测试中取得了性能飞跃,并在人工评估任务中与更大规模的前沿开源模型相媲美。

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 23)

阅读逐段中英对照全文 →