1See Contributions and Acknowledgments section for full author list. Please send correspondence to gemma4report@gmail.com.1 完整作者名单请参阅“贡献与致谢”章节。如有疑问,请发送邮件至 gemma4report@gmail.com。
Gemma 4 Technical ReportGemma 4 技术报告
Abstract摘要
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.我们推出了 Gemma 4,这是 Gemma 模型家族中新一代开源权重、原生多模态语言模型。Gemma 4 模型套件旨在提升计算效率与推理能力,包含稠密(Dense)和混合专家(MoE)架构,参数量从 23 亿到 310 亿不等。除了为所有模型尺寸改进视觉和音频编码器外,我们还为 120 亿参数模型提出了一种统一的无编码器(encoder-free)架构,可直接摄入原始音频和图像块。此外,我们集成了“思考模式”(thinking mode),使 Gemma 模型在回复前能够生成推理轨迹。通过关键的设计选择,我们提升了推理速度、内存和计算效率,以及长上下文处理能力。Gemma 4 在 STEM、多模态和长上下文基准测试中实现了性能飞跃,并在人类评测任务中可与更大规模的前沿开源模型相媲美。
1 Introduction1 引言
The rapid evolution of large language models has driven the need for open-weight models with strong multimodal understanding, reasoning, and computational efficiency. Building upon the foundations of its predecessors [Gemma Team, 2024a, b, 2025a], we introduce Gemma 4, the most capable and efficient generation in the Gemma model family to date. Gemma 4 offers natively multimodal architectures, capable of seamlessly processing text, images, and audio while achieving frontier-level performance on highly complex reasoning tasks. The Gemma 4 family is built to serve a variety of on-device hardware. The model suite includes both dense architectures (2.3B, 4.5B, 12B, and 31B parameters) and a Mixture-of-Experts [Jacobs et al., 1991, MoE] variant with 3.8B activated and 26B total parameters. We introduce several architectural and methodological innovations:大型语言模型的快速演进催生了对具备强大多模态理解、推理能力和计算效率的开源权重模型的需求。基于前代模型 [Gemma Team, 2024a, b, 2025a] 的基础,我们推出了 Gemma 4,这是 Gemma 模型家族迄今为止性能最强、效率最高的一代。Gemma 4 提供原生多模态架构,能够无缝处理文本、图像和音频,同时在高度复杂的推理任务上达到前沿水平。Gemma 4 系列专为适配各种端侧硬件而构建。模型套件既包含稠密架构(23 亿、45 亿、120 亿和 310 亿参数),也包含混合专家 [Jacobs et al., 1991, MoE] 变体(激活参数 38 亿,总参数 260 亿)。我们引入了多项架构和方法论创新:
-
•
Thinking mode for advanced reasoning: We introduce a thinking mode [OpenAI, 2024] to Gemma 4 models. By outputting a reasoning trace before the response, models demonstrate improved capabilities in reasoning-heavy domains such as mathematics and coding.用于高级推理的“思考模式”:我们为 Gemma 4 模型引入了思考模式 [OpenAI, 2024]。通过在回复前输出推理轨迹,模型在数学和编程等重推理领域展现出了更强的能力。
-
•
Long-context efficiency: Extended contexts lead to a memory explosion in the KV cache. We conserve a 5:1 ratio of local sliding window to global self-attention (4:1 for the 2.3B model) and use -RoPE [Barbero et al., 2025] as positional encoding. Combined with KV cache sharing [Shazeer, 2019] and the reuse of keys as values in global layers [Kayyam et al., 2026], these optimizations reduce the global KV cache footprint by up to 37.5%.长上下文效率:扩展上下文会导致 KV 缓存内存爆炸。我们维持了 5:1 的局部滑动窗口与全局自注意力比例(23 亿模型为 4:1),并采用 pp-RoPE [Barbero et al., 2025] 作为位置编码。结合 KV 缓存共享 [Shazeer, 2019] 以及在全局层中复用键作为值 [Kayyam et al., 2026],这些优化将全局 KV 缓存占用空间减少了高达 37.5%。
-
•
Compute efficiency: We release an autoregressive multi-token prediction (MTP) drafter head [Li et al., 2024] designed for speculative decoding [Leviathan et al., 2023] to improve the decoding speed of our models.计算效率:我们发布了一种自回归多标记预测(MTP)草稿头 [Li et al., 2024],专为推测解码 [Leviathan et al., 2023] 设计,以提高模型的解码速度。
-
•
Memory efficiency: We provide quantized versions of our models trained with quantization-aware training [Jacob et al., 2018, QAT] to reduce their parameter memory footprint and latency with minimal impact on quality.内存效率:我们提供了经过量化感知训练 [Jacob et al., 2018, QAT] 的量化版本模型,在几乎不影响质量的前提下,减少了参数内存占用和延迟。
-
•
Encoder-free architecture: Gemma 4 models have frozen vision and audio encoders. We introduce a unified encoder-free architecture for the 12B model, which projects raw 40ms audio chunks and image patches into the LLM embedding space, alleviating the need for separate encoders and reducing memory fragmentation.无编码器架构:Gemma 4 模型拥有冻结的视觉和音频编码器。我们为 120 亿模型引入了一种统一的无编码器架构,将原始 40ms 音频片段和图像块直接投影到 LLM 嵌入空间中,从而无需单独的编码器,并减少了内存碎片。
In this technical report, we outline the different model architectures across model sizes as well as the pre-training and post-training recipe of Gemma 4. Through comprehensive benchmarks and human evaluations such as Arena [Chiang et al., 2024], we demonstrate that Gemma 4 operates at a level comparable to larger, frontier open-source models across text, image, and audio modalities. We release the Gemma 4 models under an Apache 2.0 license, empowering developers and researchers everywhere to build upon, customize, and extend these capabilities.在本技术报告中,我们概述了不同模型尺寸的架构,以及 Gemma 4 的预训练和后训练方案。通过全面的基准测试和包括 Arena [Chiang et al., 2024] 在内的人类评估,我们证明了 Gemma 4 在文本、图像和音频模态上达到了与更大规模的前沿开源模型相当的水平。我们以 Apache 2.0 许可证发布 Gemma 4 模型,旨在赋能全球开发者和研究人员,基于这些能力进行构建、定制和扩展。
| Model | Audio Encoder | Vision Encoder | Embedder | Einsums | Drafter |
| E2B | 305M | 150M | 400M + 2,340M | 1,870M | 76M |
| E4B | 305M | 150M | 670M + 2,820M | 3,940M | 77M |
| 12B | - | - | 1,000M | 10,890M | 400M |
| 26B-A4B* | - | 550M | 740M | 24,500M / 2,800M (active) | 430M |
| 31B | - | 550M | 1,410M | 29,290M | 500M |
2 Model Architecture2 模型架构
Gemma 4 models follow a decoder-only Transformer architecture [Vaswani et al., 2017]. Our models have pre-norm and post-norm with RMSNorm [Zhang and Sennrich, 2019], and QKNorm [Henry et al., 2020].Gemma 4 模型遵循仅解码器(decoder-only)的 Transformer 架构 [Vaswani et al., 2017]。我们的模型采用了 RMSNorm [Zhang and Sennrich, 2019] 和 QKNorm [Henry et al., 2020] 进行前归一化和后归一化。
Dense and MoE: The Gemma 4 family of models comprises dense architectures, with effective 2.3B (E2B), effective 4.5B (E4B), 12B and 31B parameters, as well as an MoE model with 3.8B activated parameters for 26B total parameters (26B-A4B). E2B and E4B use per-layer embeddings as in Gemma 3n [Gemma Team, 2025b], making them 2.3B and 4.5B effective out of 5B and 8B total parameters respectively.稠密与 MoE:Gemma 4 系列模型包含稠密架构(有效参数 23 亿 (E2B)、有效参数 45 亿 (E4B)、120 亿和 310 亿参数),以及一个激活参数 38 亿、总参数 260 亿的 MoE 模型 (26B-A4B)。E2B 和 E4B 使用了与 Gemma 3n [Gemma Team, 2025b] 相同的逐层嵌入,使其在 50 亿和 80 亿总参数下分别实现了 23 亿和 45 亿的有效参数量。
| Shards | |||||
| Model | TPU | #Chips | Data | Seq | Replica |
| E2B | v6e | 4,096 | 16 | 8 | 32 |
| E4B | v6e | 6,144 | 16 | 16 | 24 |
| 12B | v5p | 12,288 | 16 | 16 | 48 |
| 26B-A4B* | v6e | 6,144 | 16 | 16 | 24 |
| 31B | v6e | 10,240 | 16 | 16 | 40 |
Long-context efficiency: Our local to global attention ratio patterns follow Gemma Team [2025a], that is, 4-to-1 local attention blocks for E2B and 5-to-1 for the rest. We improve memory efficiency by re-using keys as values in the global attention layers (except in E2B and E4B), i.e. , . We encode position with -RoPE with on global attention layers and with RoPE on local attention layers, effectively reducing the global KV cache by 37.5%. The RoPE frequencies are set to 1M and 10k on global and local attention layers, respectively. Finally, we share the KV cache with ratios of 20/35 and 18/42 for the E2B and E4B model.长上下文效率:我们的局部到全局注意力比例模式遵循 Gemma Team [2025a] 的方案,即 E2B 采用 4:1 的局部注意力块,其余模型采用 5:1。我们通过在全局注意力层(E2B 和 E4B 除外)中复用键作为值(即 values=keys)来提高内存效率。我们使用 pp-RoPE 对全局注意力层进行位置编码(p=0.25),并对局部注意力层使用 RoPE,有效地将全局 KV 缓存减少了 37.5%。RoPE 频率在全局和局部注意力层分别设置为 1M 和 10k。最后,我们对 E2B 和 E4B 模型分别以 20/35 和 18/42 的比例共享 KV 缓存。
2.1 Vision modality2.1 视觉模态
E2B and E4B Gemma models come with a 150M vision encoder, while larger models use a 550M encoder (except for the unified 12B). Both are Vision Transformers [Dosovitskiy et al., 2021, ViT] with a patch size of 16, whose architectural differences are detailed in Table 10 in Appendix. Our vision encoders support variable aspect ratios (see Figure 2 and Algorithm 1) and incorporate both axial 2D-RoPE [Heo et al., 2024] with non-causal attention and 2D absolute positional embeddings. We restrict the maximum number of tokens, to the values and (see Algorithm 1 for implementation details).E2B 和 E4B Gemma 模型配备了 1.5 亿参数的视觉编码器,而较大模型使用 5.5 亿参数编码器(统一的 120 亿模型除外)。两者均为视觉 Transformer [Dosovitskiy et al., 2021, ViT],补丁大小为 16,其架构差异详见附录表 10。我们的视觉编码器支持可变长宽比(见图 2 和算法 1),并结合了带有非因果注意力的轴向 2D-RoPE [Heo et al., 2024] 和 2D 绝对位置嵌入。我们将最大标记数 Nmax 限制为 70、140、280、560 和 1120(实现细节见算法 1)。
2.2 Audio modality2.2 音频模态
E2B and E4B Gemma models use a 305M audio encoder that processes audio in 40ms chunks with Mel filterbank inputs. The encoder architecture is based on the Universal Speech Model [Zhang et al., 2023, USM], consisting of two downsampling convolution layers followed by twelve Conformer layers [Gulati et al., 2020]. While the architecture remains similar to that of Gemma 3n, we reduce the number of parameters by 55% (from 680M to 305M). We do not use vector quantization; the LLM ingests the continuous representations produced by the audio encoder. As with the vision encoder, we keep weights frozen during pre-training.E2B 和 E4B Gemma 模型使用 3.05 亿参数的音频编码器,以 40ms 为单位处理音频,并以梅尔滤波器组作为输入。编码器架构基于通用语音模型 [Zhang et al., 2023, USM],由两个下采样卷积层和随后的十二个 Conformer 层 [Gulati et al., 2020] 组成。虽然架构与 Gemma 3n 保持相似,但我们将参数量减少了 55%(从 6.8 亿减至 3.05 亿)。我们不使用向量量化;LLM 直接摄入音频编码器产生的连续表示。与视觉编码器一样,我们在预训练期间保持权重冻结。
2.3 Encoder-free architecture2.3 无编码器架构
Gemma 4 12B is trained from scratch based on a new, unified, and encoder-free model paradigm, replacing the separate vision and audio encoders with lightweight projection modules. For the vision modality, Gemma 4 12B takes in 48483 RGB patches, but replaces the 550M vision encoder by a single large matmul (35M parameters). Spatial awareness is maintained by adding 2D coordinate-based positional embeddings directly to the patch representations before a final LayerNorm layer [Ba et al., 2016].Gemma 4 12B 基于一种全新的、统一的无编码器模型范式从头开始训练,用轻量级投影模块取代了单独的视觉和音频编码器。对于视觉模态,Gemma 4 12B 接收 48×48×3 的 RGB 补丁,但将 5.5 亿参数的视觉编码器替换为一个大型矩阵乘法运算(3500 万参数)。通过在最终层归一化层 [Ba et al., 2016] 之前直接向补丁表示添加基于 2D 坐标的位置嵌入,保持了空间感知能力。
For audio, the 305M USM-based conformer encoder is entirely discarded. Raw audio is segmented into 40ms chunks at 16kHz, resulting in 640-dimensional vectors per chunk. These are projected directly into the LLM embedding space. Since audio is a temporal sequence, it does not require additional positional encoding.对于音频,基于 3.05 亿参数 USM 的 Conformer 编码器被完全弃用。原始音频以 16kHz 分割为 40ms 的片段,每个片段生成 640 维向量。这些向量直接投影到 LLM 嵌入空间中。由于音频是时间序列,因此不需要额外的位置编码。
| Model | bf16 | Quantized | KV Cache |
| E2B | 4.6 | +0.05 | |
| E4B | 9.0 | +0.14 | |
| 12B | 24.0 | +0.28 | |
| 26B-A4B* | 52.0 / 7.6 | +0.28 | |
| 31B | 64.0 | +1.10 |
2.4 Pre-training2.4 预训练
We follow a similar pre-training as Gemma 3.我们遵循与 Gemma 3 类似的预训练流程。
Training data. Our pre-training dataset is a large-scale, diverse collection of data from a wide range of domains and modalities, including web documents, code, images, and audio (for E2B, E4B and 12B), with a cutoff date of January 2025.训练数据。我们的预训练数据集是一个大规模、多样化的集合,涵盖了广泛的领域和模态,包括网页文档、代码、图像和音频(针对 E2B、E4B 和 12B),截止日期为 2025 年 1 月。
Tokenizer. We use the same tokenizer as Gemini Team [2025] that is, a SentencePiece tokenizer [Kudo and Richardson, 2018] with split digits, preserved whitespace, and byte-level encodings. The vocabulary has 262k entries.分词器。我们使用与 Gemini Team [2025] 相同的分词器,即 SentencePiece 分词器 [Kudo and Richardson, 2018],支持拆分数字、保留空格和字节级编码。词表包含 26.2 万个条目。
Filtering. We filter data to decontaminate benchmarks, and to reduce the risk of unwanted or unsafe utterances and the risk of recitation.过滤。我们对数据进行过滤以去除基准测试污染,并降低有害或不安全言论以及复述风险。
| Rank | Model | Elo | 95% CI | Open | Type | #params/#activated |
| 1 | Claude Fable 5 | 1508 | 9 | no | - | - / - |
| … | ||||||
| 15 | GLM 5.1 | 1475 | 6 | yes | MoE | 744B / 40B |
| 25 | GLM 5.2 (Max) | 1471 | 10 | yes | MoE | 744B / 40B |
| 29 | MiMo V2.5 Pro | 1466 | 5 | yes | MoE | 1T / 42B |
| 34 | Kimi K2.6 | 1460 | 5 | yes | MoE | 1T / 32B |
| 36 | DeepSeek V4 Pro Thinking | 1458 | 5 | yes | MoE | 1.6T / 49B |
| 37 | GLM 5 | 1457 | 5 | yes | MoE | 744B / 40B |
| 38 | DeepSeek V4 Pro | 1456 | 5 | yes | MoE | 1.6T / 49B |
| 43 | Gemma 4 31B | 1451 | 8 | yes | Dense | 31B |
| 44 | Kimi K2.5 Thinking | 1450 | 4 | yes | MoE | 1T / 32B |
| 57 | Qwen 3.5 397B-A17B | 1444 | 4 | yes | MoE | 397B / 17B |
| 61 | Gemma 4 26B-A4B | 1438 | 8 | yes | MoE | 26B / 4B |
| 63 | DeepSeek V4 Flash Thinking | 1436 | 5 | yes | MoE | 284B / 13B |
| … | ||||||
| 157 | Gemma 3 27B | 1366 | 4 | yes | Dense | 27B |
2.5 Quantization-Aware Training2.5 量化感知训练 (QAT)
We provide quantized models and encoders in different formats along with the raw checkpoints. Based on the most popular open source quantization inference engines (e.g. llama.cpp) as well as efficient hardware support, we focus on two sets of weight representations:除了原始检查点外,我们还提供不同格式的量化模型和编码器。基于最流行的开源量化推理引擎(如 llama.cpp)以及高效的硬件支持,我们专注于两组权重表示:
-
•
mobile quantization: per-channel low bitwidth weight (mix of int2 and int4) and activation quantization (int8).移动端量化:逐通道低位宽权重(int2 和 int4 混合)和激活量化(int8)。
-
•
Q4_0 quantization: blockwise quantization, often referred to as Q4_0.Q4_0 量化:块级量化,通常称为 Q4_0。
In Table 3, we report the memory filled by raw and quantized models with and without a KV cache for a sequence of 32k tokens. Furthermore, to enable stable inference in fp16, we introduce a scalar scale at each block in order to bound the activation ranges to fit fp16.在表 3 中,我们报告了 32k 标记序列下,带有和不带有 KV 缓存的原始模型与量化模型的内存占用。此外,为了实现 fp16 下的稳定推理,我们在每个块中引入了一个标量缩放,以限制激活范围,使其适配 fp16。
| Gemma 4 | Gemma 3 | ||||||
| 31B | 26B-A4B | 12B | E4B | E2B | 27B non-thinking | ||
| MMLU Pro | 85.2 | 82.6 | 77.2 | 69.4 | 60.0 | 67.6 | |
| AIME 2026 no tools | 89.2 | 88.3 | 77.5 | 42.5 | 37.5 | 20.8 | |
| LiveCodeBench v6 | 80.0 | 77.1 | 72.0 | 52.0 | 44.0 | 29.1 | |
| Codeforces Elo | 2150 | 1718 | 1659 | 940 | 633 | 110 | |
| SciCode | 43.0 | 40.0 | 38.0 | 24.0 | 21.0 | 21.0 | |
| GPQA Diamond | 84.3 | 82.3 | 78.8 | 58.6 | 43.4 | 42.4 | |
| Big Bench Extra Hard micro avg | 74.4 | 64.8 | 53.0 | 33.1 | 21.9 | 19.3 | |
| HLE | 19.5 | 8.7 | 5.2 | - | - | - | |
| HLE with search | 26.5 | 17.2 | - | - | - | - | |
| IFBench | 76.0 | 72.0 | 74.0 | 44.0 | 38.0 | 32.0 | |
| IFEval | 98.9 | 98.5 | 97.2 | 96.7 | 94.6 | 90.4 | |
| MMMLU | 88.4 | 86.3 | 83.4 | 76.6 | 67.4 | 70.7 | |
| MRCR v2 8-needle, 128k | 66.4 | 44.1 | 43.4 | 25.4 | 19.1 | 13.5 | |
| Terminal Bench Hard | 36.0 | 14.0 | 18.0 | 8.0 | 3.0 | 4.0 | |
| Tau2 – airline | 75.0 | 76.0 | 75.0 | 52.0 | 31.0 | 39.0 | |
| Tau2 – retail | 86.4 | 85.5 | 77.6 | 67.1 | 34.6 | 6.6 | |
| Tau2 – telecom | 69.3 | 43.0 | 54.4 | 18.4 | 19.7 | 3.1 | |
We also apply QAT to the image and audio encoders. On the 150M image encoder, quantizing activations and weights to 8-bit precision (W8A8) yields a 2 reduction in total forward-pass memory footprint (from 400 MB to 200 MB, including on-device compilation overhead) and a 44% reduction in on-device latency relative to Gemma 3n on newer hardware. On the audio encoder, we further reduce activation precision to 8 bits and weight precision to bits, varying by layer cluster. Overall, we achieve a 78% reduction in on-disk footprint, from 390 MB in Gemma 3n to 87 MB in this version.我们还将 QAT 应用于图像和音频编码器。在 1.5 亿参数图像编码器上,将激活和权重量化为 8 位精度 (W8A8) 使总前向传播内存占用减少了 2 倍(从 400 MB 降至 200 MB,包括端侧编译开销),且在较新硬件上相较于 Gemma 3n 减少了 44% 的端侧延迟。在音频编码器上,我们将激活精度进一步降低至 8 位,权重精度降低至 {2,4,8} 位(根据层簇而定)。总体而言,磁盘占用空间减少了 78%,从 Gemma 3n 的 390 MB 降至本版本的 87 MB。
2.6 Multi-Token Prediction Drafter2.6 多标记预测草稿头
We train a small autoregressive MTP drafter head with our models, used for speculative decoding. In our MTP procedure, the model’s last layer activations from the previous step and token embeddings are fed into the MTP head. The MTP head generates future tokens sequentially using a separate embedder and a 4-layer Transformer block that cross-attends to the KVs of the main model (Figure 1), thus eliminating the need for MTP prefill and supporting any draft length. The Transformer block has model dimension 256 for E2B and E4B, 1024 for 26B-A4B and 31B, three local, and one global attention layers.我们随模型训练了一个小型自回归 MTP 草稿头,用于推测解码。在我们的 MTP 流程中,将模型上一层的激活值和标记嵌入输入到 MTP 头中。MTP 头使用单独的嵌入器和 4 层 Transformer 块(该块与主模型的 KV 进行交叉注意力计算,见图 1)顺序生成未来标记,从而消除了 MTP 预填充的需求并支持任意草稿长度。对于 E2B 和 E4B,Transformer 块的模型维度为 256;对于 26B-A4B 和 31B,维度为 1024,包含三个局部注意力层和一个全局注意力层。
Efficient MTP Decoding.高效 MTP 解码。
For the E2B and E4B drafters, we reduce the decoding overhead by replacing the projection operation to the entire vocabulary by a top-k operation on clusters of tokens. As a result, final matrix multiplication is reduced from to while preserving a similar acceptance rate.对于 E2B 和 E4B 草稿头,我们通过将整个词表的投影操作替换为标记簇上的 Top-k 操作来降低解码开销。因此,最终矩阵乘法从 d×262,000 减少到 d×4096,同时保持了相似的接受率。
2.7 Compute Infrastructure2.7 计算基础设施
We train our models with TPUv5p and TPUv6e as outlined in Table 2. Each model configuration is optimized to minimize training step time. For our larger models, we leverage Slice-Granularity Elasticity [Gemini Team, 2025], which allows continuous training with fewer “slices” of TPU chips when there is a localized failure. This reconfiguration reduces the delay caused by interruptions from many minutes to a few seconds.我们使用表 2 中列出的 TPUv5p 和 TPUv6e 训练模型。每个模型配置都经过优化,以最大限度地缩短训练步长。对于较大模型,我们利用切片粒度弹性 [Gemini Team, 2025],这允许在出现局部故障时使用较少的 TPU 芯片“切片”进行持续训练。这种重配置将中断造成的延迟从几分钟减少到几秒钟。
The optimizer state is sharded using an implementation of ZeRO-3 [Ren et al., 2021]. For multi-pod training, we perform a data replica reduction over the data center network, using the Pathways approach of Barham et al. [2022]. We use the single controller programming paradigm of JAX [Roberts et al., 2023] and Pathways, along with the GSPMD partitioner [Xu et al., 2021] and the MegaScale XLA compiler [XLA, 2019].优化器状态使用 ZeRO-3 [Ren et al., 2021] 的实现进行分片。对于多 POD 训练,我们使用 Barham 等人 [2022] 的 Pathways 方法在数据中心网络上执行数据副本缩减。我们使用 JAX [Roberts et al., 2023] 和 Pathways 的单控制器编程范式,以及 GSPMD 分区器 [Xu et al., 2021] 和 MegaScale XLA 编译器 [XLA, 2019]。
| Gemma 4 | Gemma 3 | ||||||
| 31B | 26B-A4B | 12B | E4B | E2B | 27B | ||
| MMMU Pro | 76.9 | 73.8 | 69.1 | 52.6 | 44.2 | 49.7 | |
| MATH-Vision | 85.6 | 82.4 | 79.7 | 59.5 | 52.4 | 46.0 | |
| MedXPertQA MM | 61.3 | 58.1 | 48.7 | 28.7 | 23.5 | - | |
| InfographicVQA | 92.0 | 89.3 | 88.4 | 70.0 | 63.9 | 70.6 | |
| OmniDocBench 1.5 | 0.131 | 0.149 | 0.164 | 0.181 | 0.290 | 0.365 | |
3 Instruction Tuning3 指令微调
Pre-trained models are turned into instruction-tuned models with a similar post-training approach as in Gemma 3. A significant difference is the addition of a thinking mode, where the model can output a reasoning trace before answering.预训练模型通过与 Gemma 3 类似的后训练方法转化为指令微调模型。一个显著的区别是增加了思考模式,模型在回答前可以输出推理轨迹。
Data filtering. We carefully optimize the data used in post-training to maximize model performance. We filter examples that show certain personal information, unsafe or toxic model outputs, mistaken self-identification data, and duplicated examples. Including subsets of data that encourage better in-context attribution, hedging, and refusals to minimize hallucinations also improves performance on factuality metrics, without degrading model performance on other metrics.数据过滤。我们仔细优化了后训练中使用的数据,以最大限度地提高模型性能。我们过滤了显示特定个人信息、不安全或有毒模型输出、错误的自我识别数据以及重复示例的条目。包含鼓励更好的上下文归因、对冲和拒绝以最小化幻觉的数据子集,也有助于提高事实性指标的性能,且不会降低模型在其他指标上的表现。
PT versus IT formatting. All models share the same tokenizer, with some control tokens dedicated to IT formatting. A key difference is that PT models output an <eos> token at the end of generation, while IT models output <turn|> at the end of the generation. An example is given for IT in Table 11. Fine-tuning either model type thus requires adding their respective end tokens. We detail how to activate thinking and how models handle function calling in Table 11.PT 与 IT 格式。所有模型共享同一个分词器,并配有专门用于 IT 格式的控制标记。一个关键区别是,PT 模型在生成结束时输出 <eos> 标记,而 IT 模型在生成结束时输出 <turn|>。表 11 给出了 IT 的示例。因此,微调任一模型类型都需要添加其各自的结束标记。我们在表 11 中详细说明了如何激活思考模式以及模型如何处理函数调用。
4 Evaluation of final models4 最终模型评估
In this section, we evaluate the IT models over a series of automated benchmarks and human evaluations across a variety of domains, as well as static benchmarks such as MMLU Pro.在本节中,我们在一系列自动化基准测试和跨领域的各种人类评估,以及 MMLU Pro 等静态基准测试上对 IT 模型进行了评估。
4.1 Human evaluation4.1 人类评估
We report the performance of our 31B and 26B-A4B models on Arena [Chiang et al., 2024] in blind side-by-side evaluations by human raters against other state-of-the-art models. We report Elo scores in Table 4. Gemma 4 31B is the top open model in the dense category, and both Gemma 4 31B and 26B-A4B show performance equal to much larger open models.我们报告了 31B 和 26B-A4B 模型在 Arena [Chiang et al., 2024] 上的表现,该评估通过人类评估员进行的盲测双盲评估与其他最先进模型进行对比。我们在表 4 中报告了 Elo 分数。Gemma 4 31B 是稠密类别中的顶级开源模型,Gemma 4 31B 和 26B-A4B 的表现均等同于更大规模的开源模型。
| CoVoST (CorpusBLEU ) | ||||||||||
| Params | Size | ja en | de en | fr en | es en | it en | ru en | zh en | AVG | |
| Gemma 4 E2B | 305M | 87 MB | 21.4 | 39.2 | 39.2 | 43.2 | 40.8 | 46.4 | 17.9 | 35.4 |
| Gemma 4 E4B | 25.5 | 42.0 | 41.0 | 44.8 | 43.0 | 49.4 | 21.9 | 38.2 | ||
| Gemma 3n E2B | 680M | 390 MB | 17.7 | 36.5 | 35.7 | 39.9 | 38.5 | 39.2 | 13.9 | 31.6 |
| Gemma 3n E4B | 22.3 | 39.1 | 38.4 | 41.8 | 40.4 | 43.7 | 17.4 | 34.7 | ||
| FLEURS ASR (WER , * = CER ) | |||||||||||||
| en | ko* | ja* | de | fr | hi | es | it | pt-br | ru | ar | zh* | AVG | |
| Gemma 4 E2B | 0.080 | 0.066 | 0.107 | 0.076 | 0.101 | 0.101 | 0.042 | 0.041 | 0.056 | 0.084 | 0.143 | 0.187 | 0.090 |
| Gemma 4 E4B | 0.065 | 0.053 | 0.078 | 0.061 | 0.080 | 0.086 | 0.035 | 0.032 | 0.046 | 0.068 | 0.162 | 0.136 | 0.075 |
| Gemma 3n E2B | 0.076 | 0.101 | 0.163 | 0.079 | 0.130 | 0.106 | 0.051 | 0.044 | 0.067 | 0.112 | 0.131 | 0.235 | 0.108 |
| Gemma 3n E4B | 0.066 | 0.073 | 0.111 | 0.065 | 0.098 | 0.089 | 0.041 | 0.034 | 0.053 | 0.087 | 0.101 | 0.203 | 0.085 |
4.2 Static benchmarks4.2 静态基准测试
In Table 5, we show the performance of our final models across a variety of benchmarks compared to Gemma 3 27B. Gemma 4 31B is closest in size and significantly better across the board, while E2B roughly matches Gemma 3 27B performance with 10x less parameters. Table 6 shows the performance of Gemma 4 models on vision benchmarks, with E4B equaling or outperforming Gemma 3 27B on all evals. Tables 7 and 8 display the multilingual audio transcription and translation performance of E2B & E4B and of 12B respectively. Table 9 shows a leap on long-context capabilities between Gemma 3 27B and Gemma 4 models, with E4B outperforming Gemma 3 27B.在表 5 中,我们展示了最终模型在各种基准测试上与 Gemma 3 27B 的对比表现。Gemma 4 31B 在尺寸上最接近,且全面显著优于后者,而 E2B 仅用十分之一的参数就大致匹配了 Gemma 3 27B 的性能。表 6 展示了 Gemma 4 模型在视觉基准测试上的表现,E4B 在所有评估中均持平或优于 Gemma 3 27B。表 7 和表 8 分别展示了 E2B & E4B 以及 12B 的多语言音频转录和翻译性能。表 9 展示了 Gemma 3 27B 与 Gemma 4 模型在长上下文能力上的飞跃,E4B 优于 Gemma 3 27B。
5 Responsibility, Safety, Security5 责任、安全与保障
As open models become central to enterprise infrastructure, provenance and security are paramount. Gemma 4 undergoes the same rigorous safety evaluations as Gemini models. Responsibility, safety, and security are of utmost importance in the development workflow, ensuring that these language models are designed from the ground up for responsible AI development.随着开源模型成为企业基础设施的核心,来源和安全性至关重要。Gemma 4 经历了与 Gemini 模型相同的严格安全评估。责任、安全和保障在开发工作流程中至关重要,确保这些语言模型从设计之初就以负责任的 AI 开发为目标。
5.1 Governance & Assessment5.1 治理与评估
Our approach to assessing the benefits and risks of Gemma 4 reflects the foundation established in prior models, updated to account for its expanded multimodal capabilities. We maintain the belief that openness in AI can spread the benefits of these technologies across society, but this must be continuously evaluated against the risk of malicious uses that can cause individual and institutional harm [Weidinger et al., 2021].我们评估 Gemma 4 收益与风险的方法反映了前代模型建立的基础,并根据其扩展的多模态能力进行了更新。我们始终认为,AI 的开放性可以将这些技术的益处传播给整个社会,但这必须根据可能导致个人和机构伤害的恶意使用风险进行持续评估 [Weidinger et al., 2021]。
Gemma 4 models were developed in partnership with internal safety and responsible AI teams. Releasing these models required careful scrutiny of the evolving risks associated with LLMs and an understanding of how models are deployed in the wild. While an open model shares innovation across the AI ecosystem, we remain committed to providing educational resources to users and monitoring downstream model usage.Gemma 4 模型是与内部安全和负责任 AI 团队合作开发的。发布这些模型需要仔细审查与 LLM 相关的演变风险,并了解模型在现实世界中的部署方式。虽然开源模型在 AI 生态系统中共享创新,但我们仍致力于为用户提供教育资源并监控下游模型使用情况。
| FLEURS ASR (WER , * = CER ) | ||||
| en | ko* | ja* | de | fr |
| 0.063 | 0.057 | 0.080 | 0.053 | 0.081 |
| es | it | pt-br | ru | ar |
| 0.038 | 0.030 | 0.047 | 0.068 | 0.070 |
| CoVoST (XX EN, CorpusBLEU ) | |||||
| ja | de | fr | es | it | ru |
| 26.4 | 41.9 | 42.5 | 44.6 | 43.3 | 50.5 |
| Gemma 4 | Gemma 3 | ||||||||||
| Benchmark | Metric | Context length | 31B | 26B-A4B | 12B | E4B | E2B | 27B | |||
| RULER | Accuracy | 32k | 96.8 | 97.3 | 96.4 | 95.2 | 83.0 | 91.1 | |||
| 128k | 96.4 | 89.8 | 91.2 | 86.6 | 70.4 | 66.0 | |||||
|
Recall@k | 128k | 79.5 | 66.3 | 66.4 | 58.5 | 50.5 | 8.6 | |||
| GraphWalks | F1 | <128k | 82.3 | 72.6 | 71.0 | 50.9 | 4.1 | 32.8 | |||
| MTOB | chrF | 128k (Half book) | 52.9 | 50.0 | 45.1 | 37.8 | 15.4 | 41.0 | |||
| (engkgv) | 256k (Full book) | 54.3 | 48.9 | 41.9 | - | - | - | ||||
| MTOB | chrF | 128k (Half book) | 48.6 | 45.0 | 37.3 | 34.6 | 28.2 | 31.2 | |||
| (kgveng) | 256k (Full book) | 46.2 | 42.7 | 32.9 | - | - | - | ||||
5.2 Safety Policies and Train-Time Mitigations5.2 安全政策与训练时缓解措施
A key pillar of Gemma’s safety approach is aligning our fine-tuned models with Google’s AI principles and safety policies. These policies aim to prevent our generative models from producing harmful content, specifically:Gemma 安全方法的一个关键支柱是将微调模型与 Google 的 AI 原则和安全政策保持一致。这些政策旨在防止我们的生成模型产生有害内容,具体包括:
-
•
Content related to child sexual abuse material (CSAM) and exploitation;与儿童性虐待材料 (CSAM) 和剥削相关的内容;
-
•
Dangerous content, e.g., promoting suicide, or instructing in activities that could cause real-world harm;危险内容,例如宣扬自杀或指导可能导致现实世界伤害的活动;
-
•
Sexually explicit content;色情内容;
-
•
Hate speech, e.g., dehumanizing members of protected groups;仇恨言论,例如对受保护群体成员进行非人化处理;
-
•
Harassment, e.g., encouraging violence against people.骚扰,例如鼓励针对个人的暴力。
To mitigate these risks, Gemma 4 models underwent careful input data pre-processing and scrutiny. The training data was specifically filtered for the removal of certain personal information and other sensitive data to guard against privacy violations. Post-training evaluations and train-time mitigations were also implemented to align the model with our safety policies.为了减轻这些风险,Gemma 4 模型经过了仔细的输入数据预处理和审查。训练数据经过专门过滤,删除了某些个人信息和其他敏感数据,以防止隐私泄露。还实施了后训练评估和训练时缓解措施,以使模型符合我们的安全政策。
5.3 Safety Evaluations5.3 安全评估
We conduct rigorous automated and human evaluations to understand the potential harms our models might cause. For all areas of safety testing, we saw major improvements in every category of content safety relative to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 and 3n models in improving safety, while keeping unjustified refusals low.我们进行严格的自动化和人类评估,以了解模型可能造成的潜在危害。在所有安全测试领域,我们看到内容安全在每个类别中相对于之前的 Gemma 模型都有重大改进。总体而言,Gemma 4 模型在提高安全性方面显著优于 Gemma 3 和 3n 模型,同时保持了较低的不正当拒绝率。
Importantly, all testing was conducted without safety filters to accurately evaluate the model’s inherent capabilities and behaviors. For both text-to-text and image-to-text modalities, and across all model sizes, the models produced minimal policy violations. We balance development speed with targeted safety testing, upholding the commitments laid out in our Frontier Safety Framework [Google DeepMind, 2024].重要的是,所有测试均在没有安全过滤器的情况下进行,以准确评估模型的内在能力和行为。对于文本到文本和图像到文本模态,以及所有模型尺寸,模型产生的政策违规极少。我们在开发速度与有针对性的安全测试之间取得平衡,秉承我们在《前沿安全框架》[Google DeepMind, 2024] 中做出的承诺。
5.4 Ethical Considerations and Risk Mitigation5.4 伦理考量与风险缓解
The development of LLMs introduces specific ethical considerations. In making Gemma 4, we focused heavily on:LLM 的开发引入了特定的伦理考量。在制作 Gemma 4 时,我们重点关注:
-
•
Bias and Fairness: LLMs trained on large-scale text and image data can reflect embedded socio-cultural biases. We encourage developers to perform continuous monitoring (using evaluation metrics and human review) and explore de-biasing techniques during model fine-tuning.偏见与公平:在海量文本和图像数据上训练的 LLM 可能反映嵌入的社会文化偏见。我们鼓励开发者进行持续监控(使用评估指标和人工审查),并在模型微调期间探索去偏技术。
-
•
Misinformation and Misuse: LLMs can be misused to generate false or misleading text. We provide technical limitations, developer education, and guidelines for responsible use within the Responsible Generative AI Toolkit to mitigate malicious applications.错误信息与滥用:LLM 可能被滥用于生成虚假或误导性文本。我们在《负责任生成式 AI 工具包》中提供了技术限制、开发者教育和负责任使用指南,以减轻恶意应用。
-
•
Privacy Considerations: While our training datasets were filtered to remove certain personal information and other sensitive data, developers are strongly encouraged to adhere to local privacy regulations and implement privacy-preserving techniques in their applications.隐私考量:虽然我们的训练数据集经过过滤以删除某些个人信息和其他敏感数据,但我们强烈建议开发者遵守当地隐私法规,并在其应用中实施隐私保护技术。
5.5 Our Approach to Responsible Open Models5.5 我们对负责任开源模型的方法
Designing safe, secure, and responsible applications requires a system-level approach that mitigates risks associated with specific use cases and environments. We provide guidelines, mechanisms, and safeguards for content safety, and encourage developers to implement appropriate configurations based on their product policies. We will continue to adopt safety mitigations proportionate to potential risks, sharing these models with the community only when confident that the benefits significantly outweigh foreseeable risks.设计安全、可靠和负责任的应用需要系统级的方法,以减轻与特定用例和环境相关的风险。我们提供内容安全的指南、机制和保障措施,并鼓励开发者根据其产品政策实施适当的配置。我们将继续采取与潜在风险相称的安全缓解措施,仅在确信收益显著大于可预见风险时,才与社区共享这些模型。
6 Discussion and Conclusion6 讨论与结论
In this technical report, we presented Gemma 4, an open-weight model family featuring multimodal dense and MoE architectures designed for varied hardware environments. Gemma 4 models come with a thinking mode in which they generate reasoning traces prior to responding, improving overall performance. We introduced a unified, encoder-free architecture that processes raw audio and image patches. We also alleviated long-context memory limitations via better local-to-global attention ratios, positional encoding, and KV cache sharing. We increased the overall compute efficiency via QAT and memory efficiency via MTP drafters. Gemma 4 models demonstrate a leap in performance compared to Gemma 3 across benchmarks, and human evaluations demonstrate that Gemma 4 performs comparably to significantly larger open models, providing a scalable foundation for edge deployment and reasoning while supporting open research.在本技术报告中,我们介绍了 Gemma 4,这是一个开源权重模型家族,具有针对不同硬件环境设计的多模态稠密和 MoE 架构。Gemma 4 模型配备了思考模式,在回答前生成推理轨迹,从而提升了整体性能。我们引入了一种处理原始音频和图像块的统一无编码器架构。我们还通过改进局部到全局注意力比例、位置编码和 KV 缓存共享,缓解了长上下文内存限制。我们通过 QAT 提高了整体计算效率,并通过 MTP 草稿头提高了内存效率。Gemma 4 模型在基准测试中展现出相较于 Gemma 3 的性能飞跃,人类评估表明 Gemma 4 的表现可与规模大得多的开源模型媲美,为边缘部署和推理提供了可扩展的基础,同时支持开放研究。
References参考文献
- Ba et al. [2016] J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. Ba et al. [2016] J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
- Barbero et al. [2025] F. Barbero, A. Vitvitskyi, C. Perivolaropoulos, R. Pascanu, and P. Veličković. Round and round we go! what makes rotary positional encodings useful? In The Thirteenth International Conference on Learning Representations, 2025. Barbero et al. [2025] F. Barbero, A. Vitvitskyi, C. Perivolaropoulos, R. Pascanu, and P. Veličković. Round and round we go! what makes rotary positional encodings useful? In The Thirteenth International Conference on Learning Representations, 2025.
- Barham et al. [2022] P. Barham, A. Chowdhery, J. Dean, S. Ghemawat, S. Hand, D. Hurt, M. Isard, H. Lim, R. Pang, S. Roy, B. Saeta, P. Schuh, R. Sepassi, L. E. Shafey, C. A. Thekkath, and Y. Wu. Pathways: Asynchronous distributed dataflow for ml, 2022. Barham et al. [2022] P. Barham, A. Chowdhery, J. Dean, S. Ghemawat, S. Hand, D. Hurt, M. Isard, H. Lim, R. Pang, S. Roy, B. Saeta, P. Schuh, R. Sepassi, L. E. Shafey, C. A. Thekkath, and Y. Wu. Pathways: Asynchronous distributed dataflow for ml, 2022.
- Barres et al. [2025] V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan. -bench: Evaluating conversational agents in a dual-control environment, 2025. Barres et al. [2025] V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan. τ2-bench: Evaluating conversational agents in a dual-control environment, 2025.
- Chiang et al. [2024] W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024. Chiang et al. [2024] W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024.
- Conneau et al. [2023] A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna. Fleurs: Few-shot learning evaluation of universal representations of speech. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805. IEEE, 2023. Conneau et al. [2023] A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna. Fleurs: Few-shot learning evaluation of universal representations of speech. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805. IEEE, 2023.
- Dosovitskiy et al. [2021] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. Dosovitskiy et al. [2021] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
- for AI Safety et al. [2026] C. for AI Safety et al. A benchmark of expert-level academic questions to assess ai capabilities. Nature, 649(8099):1139–1146, 2026. for AI Safety et al. [2026] C. for AI Safety et al. A benchmark of expert-level academic questions to assess ai capabilities. Nature, 649(8099):1139–1146, 2026.
- Gemini Team [2025] Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. Gemini Team [2025] Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
- Gemma Team [2024a] Gemma Team. Gemma: Open models based on gemini research and technology, 2024a. Gemma Team [2024a] Gemma Team. Gemma: Open models based on gemini research and technology, 2024a.
- Gemma Team [2024b] Gemma Team. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024b. Gemma Team [2024b] Gemma Team. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024b.
- Gemma Team [2025a] Gemma Team. Gemma 3: Technical report. arXiv preprint arXiv:2503.19786, 2025a. Gemma Team [2025a] Gemma Team. Gemma 3: Technical report. arXiv preprint arXiv:2503.19786, 2025a.
- Gemma Team [2025b] Gemma Team. Gemma 3n. https://deepmind.google/models/gemma/gemma-3n/, 2025b. Gemma Team [2025b] Gemma Team. Gemma 3n. https://deepmind.google/models/gemma/gemma-3n/, 2025b.
- Google DeepMind [2024] Google DeepMind. Introducing the frontier safety framework. https://deepmind.google/blog/introducing-the-frontier-safety-framework/, 2024. Google DeepMind [2024] Google DeepMind. Introducing the frontier safety framework. https://deepmind.google/blog/introducing-the-frontier-safety-framework/, 2024.
- Gulati et al. [2020] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100, 2020. Gulati et al. [2020] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100, 2020.
- Henry et al. [2020] A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4246–4253, 2020. Henry et al. [2020] A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4246–4253, 2020.
- Heo et al. [2024] B. Heo, S. Park, D. Han, and S. Yun. Rotary position embedding for vision transformer. In European Conference on Computer Vision, pages 289–305. Springer, 2024. Heo et al. [2024] B. Heo, S. Park, D. Han, and S. Yun. Rotary position embedding for vision transformer. In European Conference on Computer Vision, pages 289–305. Springer, 2024.
- Hsieh et al. [2024] C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models? arXiv preprint arXiv:2404.06654, 2024. Hsieh et al. [2024] C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models? arXiv preprint arXiv:2404.06654, 2024.
- Jacob et al. [2018] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In CVPR, 2018. Jacob et al. [2018] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In CVPR, 2018.
- Jacobs et al. [1991] R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3:79–87, 1991. Jacobs et al. [1991] R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3:79–87, 1991.
- Jain et al. [2025] N. Jain, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In International Conference on Learning Representations, volume 2025, pages 58791–58831, 2025. Jain et al. [2025] N. Jain, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In International Conference on Learning Representations, volume 2025, pages 58791–58831, 2025.
- Kayyam et al. [2026] A. Kayyam, A. M. Gopal, and M. A. Lewis. Do transformers need three projections? systematic study of qkv variants. arXiv preprint arXiv:2606.04032, 2026. Kayyam et al. [2026] A. Kayyam, A. M. Gopal, and M. A. Lewis. Do transformers need three projections? systematic study of qkv variants. arXiv preprint arXiv:2606.04032, 2026.
- Kazemi et al. [2025] M. Kazemi, B. Fatemi, H. Bansal, J. Palowitch, C. Anastasiou, S. V. Mehta, L. K. Jain, V. Aglietti, D. Jindal, P. Chen, et al. Big-bench extra hard. arXiv preprint arXiv:2502.19187, 2025. Kazemi et al. [2025] M. Kazemi, B. Fatemi, H. Bansal, J. Palowitch, C. Anastasiou, S. V. Mehta, L. K. Jain, V. Aglietti, D. Jindal, P. Chen, et al. Big-bench extra hard. arXiv preprint arXiv:2502.19187, 2025.
- Kudo and Richardson [2018] T. Kudo and J. Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. 2018. Kudo and Richardson [2018] T. Kudo and J. Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. 2018.
- Lee et al. [2024] J. Lee, A. Chen, Z. Dai, D. Dua, D. S. Sachan, M. Boratko, Y. Luan, S. M. R. Arnold, V. Perot, S. Dalmia, H. Hu, X. Lin, P. Pasupat, A. Amini, J. R. Cole, S. Riedel, I. Naim, M.-W. Chang, and K. Guu. Can long-context language models subsume retrieval, rag, sql, and more? ArXiv, 2024. Lee et al. [2024] J. Lee, A. Chen, Z. Dai, D. Dua, D. S. Sachan, M. Boratko, Y. Luan, S. M. R. Arnold, V. Perot, S. Dalmia, H. Hu, X. Lin, P. Pasupat, A. Amini, J. R. Cole, S. Riedel, I. Naim, M.-W. Chang, and K. Guu. Can long-context language models subsume retrieval, rag, sql, and more? ArXiv, 2024.
- Leviathan et al. [2023] Y. Leviathan, M. Kalman, and Y. Matias. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023. Leviathan et al. [2023] Y. Leviathan, M. Kalman, and Y. Matias. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
- Li et al. [2024] Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE: Speculative sampling requires rethinking feature uncertainty. In International Conference on Machine Learning, 2024. Li et al. [2024] Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE: Speculative sampling requires rethinking feature uncertainty. In International Conference on Machine Learning, 2024.
- Mathew et al. [2022] M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar. Infographicvqa. In WACV, 2022. Mathew et al. [2022] M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar. Infographicvqa. In WACV, 2022.
- Merrill et al. [2026] M. A. Merrill, A. G. Shaw, N. Carlini, B. Li, H. Raj, I. Bercovich, L. Shi, J. Y. Shin, T. Walshe, E. K. Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868, 2026. Merrill et al. [2026] M. A. Merrill, A. G. Shaw, N. Carlini, B. Li, H. Raj, I. Bercovich, L. Shi, J. Y. Shin, T. Walshe, E. K. Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868, 2026.
- OpenAI [2024] OpenAI. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024. OpenAI [2024] OpenAI. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
- OpenAI [2025] OpenAI. GraphWalks dataset, 2025. OpenAI [2025] OpenAI. GraphWalks dataset, 2025.
- Ouyang et al. [2025] L. Ouyang, Y. Qu, H. Zhou, J. Zhu, R. Zhang, Q. Lin, B. Wang, Z. Zhao, M. Jiang, X. Zhao, et al. Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24838–24848, 2025. Ouyang et al. [2025] L. Ouyang, Y. Qu, H. Zhou, J. Zhu, R. Zhang, Q. Lin, B. Wang, Z. Zhao, M. Jiang, X. Zhao, et al. Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24838–24848, 2025.
- Phan et al. [2025] L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. Humanity’s last exam. arXiv preprint arXiv:2501.14249, 2025. Phan et al. [2025] L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. Humanity’s last exam. arXiv preprint arXiv:2501.14249, 2025.
- Pyatkin et al. [2026] V. Pyatkin, S. Malik, V. Graf, H. Ivison, S. Huang, P. Dasigi, N. Lambert, and H. Hajishirzi. Generalizing verifiable instruction following. Advances in Neural Information Processing Systems, 38, 2026. Pyatkin et al. [2026] V. Pyatkin, S. Malik, V. Graf, H. Ivison, S. Huang, P. Dasigi, N. Lambert, and H. Hajishirzi. Generalizing verifiable instruction following. Advances in Neural Information Processing Systems, 38, 2026.
- Rein et al. [2023] D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark. ArXiv, abs/2311.12022, 2023. Rein et al. [2023] D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark. ArXiv, abs/2311.12022, 2023.
- Ren et al. [2021] J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He. Zero-offload: Democratizing billion-scale model training. In USENIX, 2021. Ren et al. [2021] J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He. Zero-offload: Democratizing billion-scale model training. In USENIX, 2021.
- Roberts et al. [2023] A. Roberts, H. W. Chung, G. Mishra, A. Levskaya, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, et al. Scaling up models and data with t5x and seqio. JMLR, 2023. Roberts et al. [2023] A. Roberts, H. W. Chung, G. Mishra, A. Levskaya, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, et al. Scaling up models and data with t5x and seqio. JMLR, 2023.
- Shazeer [2019] N. Shazeer. Fast transformer decoding: One write-head is all you need. CoRR, abs/1911.02150, 2019. Shazeer [2019] N. Shazeer. Fast transformer decoding: One write-head is all you need. CoRR, abs/1911.02150, 2019.
- Tanzer et al. [2024] G. Tanzer, M. Suzgun, E. Visser, D. Jurafsky, and L. Melas-Kyriazi. A benchmark for learning to translate a new language from one grammar book. In The Twelfth International Conference on Learning Representations, 2024. Tanzer et al. [2024] G. Tanzer, M. Suzgun, E. Visser, D. Jurafsky, and L. Melas-Kyriazi. A benchmark for learning to translate a new language from one grammar book. In The Twelfth International Conference on Learning Representations, 2024.
- Team et al. [2026] K. Team, T. Bai, Y. Bai, Y. Bao, S. Cai, Y. Cao, Y. Charles, H. Che, C. Chen, G. Chen, et al. Kimi k2. 5: Visual agentic intelligence. arXiv preprint arXiv:2602.02276, 2026. Team et al. [2026] K. Team, T. Bai, Y. Bai, Y. Bao, S. Cai, Y. Cao, Y. Charles, H. Che, C. Chen, G. Chen, et al. Kimi k2. 5: Visual agentic intelligence. arXiv preprint arXiv:2602.02276, 2026.
- Team [2026] Q. Team. Qwen3. 5-omni technical report. arXiv preprint arXiv:2604.15804, 2026. Team [2026] Q. Team. Qwen3. 5-omni technical report. arXiv preprint arXiv:2604.15804, 2026.
- Tian et al. [2024] M. Tian, L. Gao, S. D. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y. Li, et al. Scicode: A research coding benchmark curated by scientists. Advances in Neural Information Processing Systems, 37:30624–30650, 2024. Tian 等人 [2024] M. Tian, L. Gao, S. D. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y. Li 等人。 Scicode:一个由科学家策划的研究编码基准测试。 神经信息处理系统进展 (Advances in Neural Information Processing Systems), 37:30624–30650, 2024。
- Vaswani et al. [2017] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. 2017. Vaswani 等人 [2017] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, 和 I. Polosukhin。 Attention is all you need (注意力机制即你所需)。 2017。
- Vodrahalli et al. [2024] K. Vodrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi, et al. Michelangelo: Long context evaluations beyond haystacks via latent structure queries. arXiv preprint arXiv:2409.12640, 2024. Vodrahalli 等人 [2024] K. Vodrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi 等人。 Michelangelo:通过潜在结构查询实现超越大海捞针的长上下文评估。 arXiv 预印本 arXiv:2409.12640, 2024。
- Wang et al. [2020] C. Wang, J. Pino, A. Wu, and J. Gu. Covost: A diverse multilingual speech-to-text translation corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4197–4203, 2020. Wang 等人 [2020] C. Wang, J. Pino, A. Wu, 和 J. Gu。 Covost:一个多样化的多语言语音转文本翻译语料库。 收录于:第十二届语言资源与评估会议论文集,第 4197–4203 页,2020。
- Wang et al. [2024a] K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, and H. Li. Measuring multimodal mathematical reasoning with math-vision dataset. Advances in Neural Information Processing Systems, 37:95095–95169, 2024a. Wang 等人 [2024a] K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, 和 H. Li。 使用 math-vision 数据集衡量多模态数学推理。 神经信息处理系统进展 (Advances in Neural Information Processing Systems), 37:95095–95169, 2024a。
- Wang et al. [2024b] Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In NeurIPS, 2024b. Wang 等人 [2024b] Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang 等人。 Mmlu-pro:一个更稳健且更具挑战性的多任务语言理解基准测试。 收录于:NeurIPS, 2024b。
- Weidinger et al. [2021] L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel. Ethical and social risks of harm from language models, 2021. Weidinger 等人 [2021] L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, 和 I. Gabriel。 语言模型带来的伦理和社会危害风险,2021。
- Xiao et al. [2026] B. Xiao, B. Xia, B. Yang, B. Gao, B. Shen, C. Zhang, C. He, C. Lou, F. Luo, G. Wang, et al. Mimo-v2-flash technical report. arXiv preprint arXiv:2601.02780, 2026. Xiao 等人 [2026] B. Xiao, B. Xia, B. Yang, B. Gao, B. Shen, C. Zhang, C. He, C. Lou, F. Luo, G. Wang 等人。 Mimo-v2-flash 技术报告。 arXiv 预印本 arXiv:2601.02780, 2026。
- XLA [2019] XLA. Xla: Optimizing compiler for tensorflow, 2019. XLA [2019] XLA。 Xla:Tensorflow 的优化编译器,2019。
- Xu et al. [2026] A. Xu, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Ling, et al. Deepseek-v4: Towards highly efficient million-token context intelligence. arXiv preprint arXiv:2606.19348, 2026. Xu 等人 [2026] A. Xu, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Ling 等人。 Deepseek-v4:迈向高效的百万 token 上下文智能。 arXiv 预印本 arXiv:2606.19348, 2026。
- Xu et al. [2021] Y. Xu, H. Lee, D. Chen, B. A. Hechtman, Y. Huang, R. Joshi, M. Krikun, D. Lepikhin, A. Ly, M. Maggioni, R. Pang, N. Shazeer, S. Wang, T. Wang, Y. Wu, and Z. Chen. GSPMD: general and scalable parallelization for ML computation graphs. 2021. Xu 等人 [2021] Y. Xu, H. Lee, D. Chen, B. A. Hechtman, Y. Huang, R. Joshi, M. Krikun, D. Lepikhin, A. Ly, M. Maggioni, R. Pang, N. Shazeer, S. Wang, T. Wang, Y. Wu, 和 Z. Chen。 GSPMD:面向机器学习计算图的通用且可扩展的并行化方案。 2021。
- Yue et al. [2025] X. Yue, T. Zheng, Y. Ni, Y. Wang, K. Zhang, S. Tong, Y. Sun, B. Yu, G. Zhang, H. Sun, et al. Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15134–15186, 2025. Yue 等人 [2025] X. Yue, T. Zheng, Y. Ni, Y. Wang, K. Zhang, S. Tong, Y. Sun, B. Yu, G. Zhang, H. Sun 等人。 Mmmu-pro:一个更稳健的多学科多模态理解基准测试。 收录于:第 63 届计算语言学协会年会论文集(第 1 卷:长论文),第 15134–15186 页,2025。
- Zeng et al. [2026] A. Zeng, X. Lv, Z. Hou, Z. Du, Q. Zheng, B. Chen, D. Yin, C. Ge, C. Huang, C. Xie, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. Zeng 等人 [2026] A. Zeng, X. Lv, Z. Hou, Z. Du, Q. Zheng, B. Chen, D. Yin, C. Ge, C. Huang, C. Xie 等人。 Glm-5:从 vibe 编码到智能体工程。 arXiv 预印本 arXiv:2602.15763, 2026。
- Zhang and Sennrich [2019] B. Zhang and R. Sennrich. Root mean square layer normalization. 2019. Zhang 和 Sennrich [2019] B. Zhang 和 R. Sennrich。 均方根层归一化 (Root mean square layer normalization)。 2019。
- Zhang et al. [2023] Y. Zhang, W. Han, J. Qin, Y. Wang, A. Bapna, Z. Chen, N. Chen, B. Li, V. Axelrod, G. Wang, et al. Google usm: Scaling automatic speech recognition beyond 100 languages. arXiv preprint arXiv:2303.01037, 2023. Zhang 等人 [2023] Y. Zhang, W. Han, J. Qin, Y. Wang, A. Bapna, Z. Chen, N. Chen, B. Li, V. Axelrod, G. Wang 等人。 Google usm:将自动语音识别扩展到 100 多种语言。 arXiv 预印本 arXiv:2303.01037, 2023。
- Zhou et al. [2023] J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023. Zhou 等人 [2023] J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, 和 L. Hou。 大型语言模型的指令遵循评估。 arXiv 预印本 arXiv:2311.07911, 2023。
- Zuo et al. [2025] Y. Zuo, S. Qu, Y. Li, Z. Chen, X. Zhu, E. Hua, K. Zhang, N. Ding, and B. Zhou. Medxpertqa: Benchmarking expert-level medical reasoning and understanding. arXiv preprint arXiv:2501.18362, 2025. Zuo 等人 [2025] Y. Zuo, S. Qu, Y. Li, Z. Chen, X. Zhu, E. Hua, K. Zhang, N. Ding, 和 B. Zhou。 Medxpertqa:对专家级医学推理和理解能力的基准测试。 arXiv 预印本 arXiv:2501.18362, 2025。
Core contributors
核心贡献者
Sherif El Abd
Vaibhav Aggarwal
Robin Algayres
Alek Andreev
Olivier Bachem
Ian Ballantyne
Cormac Brick
Victor Cărbune
Michelle Casbon
Mayank Chaturvedi
Victor Cotruta
Alice Coucke
Phil Culliton
Robert Dadashi
Lucas Dixon
Mohamed Elhawaty
Utku Evci
Clément Farabet
Johan Ferret
Filippo Galgani
Sertan Girgin
Jean-Bastien Grill
Maarten Grootendorst
Jiaxian Guo
Cassidy Hardin
Yanzhang He
Steven M. Hernandez
Omri Homburger
Léonard Hussenot
Juyeong Ji
Armand Joulin
Aishwarya Kamath
Parnian Kassraie
Olivier Lacombe
Preethi Lahoti
Gaël Liu
Gus Martins
Luciano Martins
Tatiana Matejovicova
Ramona Merhej
Nikola Momchev
Sneha Mondal
Ryan Mullins
Sindhu Raghuram Panyam
Shreya Pathak
Sarah Perrin
André Susano Pinto
Etienne Pot
Angéline Pouget
Alexandre Ramé
Sabela Ramos
Douglas Reid
David Rim
Morgane Rivière
Karsten Roth
Louis Rouillard
Omar Sanseviero
Pier Giuseppe Sessa
Shane Settle
Danila Sinopalnikov
Sara Smoot
Piotr Stanczyk
Andreas Steiner
Lawrence Stewart
Ilya Tolstikhin
Michael Tschannen
Anton Tsitsulin
Nino Vieillard
Renjie Wu
Pingmei Xu
Haichuan Yang
Edouard Yvinec
Li Zhang
Joe Zou
Sherif El Abd
Vaibhav Aggarwal
Robin Algayres
Alek Andreev
Olivier Bachem
Ian Ballantyne
Cormac Brick
Victor Cărbune
Michelle Casbon
Mayank Chaturvedi
Victor Cotruta
Alice Coucke
Phil Culliton
Robert Dadashi
Lucas Dixon
Mohamed Elhawaty
Utku Evci
Clément Farabet
Johan Ferret
Filippo Galgani
Sertan Girgin
Jean-Bastien Grill
Maarten Grootendorst
Jiaxian Guo
Cassidy Hardin
Yanzhang He
Steven M. Hernandez
Omri Homburger
Léonard Hussenot
Juyeong Ji
Armand Joulin
Aishwarya Kamath
Parnian Kassraie
Olivier Lacombe
Preethi Lahoti
Gaël Liu
Gus Martins
Luciano Martins
Tatiana Matejovicova
Ramona Merhej
Nikola Momchev
Sneha Mondal
Ryan Mullins
Sindhu Raghuram Panyam
Shreya Pathak
Sarah Perrin
André Susano Pinto
Etienne Pot
Angéline Pouget
Alexandre Ramé
Sabela Ramos
Douglas Reid
David Rim
Morgane Rivière
Karsten Roth
Louis Rouillard
Omar Sanseviero
Pier Giuseppe Sessa
Shane Settle
Danila Sinopalnikov
Sara Smoot
Piotr Stanczyk
Andreas Steiner
Lawrence Stewart
Ilya Tolstikhin
Michael Tschannen
Anton Tsitsulin
Nino Vieillard
Renjie Wu
Pingmei Xu
Haichuan Yang
Edouard Yvinec
Li Zhang
Joe Zou
Contributors
贡献者
Nicolas Aagnes
Abdelrahman Abdelhamed
Shivani Agrawal
Shubham Agrawal
Ibrahim Alabdulmohsin
Jean Baptiste Alayrac
Uri Alon
Chandramouli Amarnath
Ankesh Anand
Chrysovalantis Anastasiou
Setareh Ariafar
François-Xavier Aubet
Kyriakos Axiotis
Federico Barbero
Joelle Barral
Alexei Bendebury
Urs Bergmann
Stanley Bileschi
Kat Black
Mathieu Blondel
Sebastian Borgeaud
Arthur Bražinskas
Ryan Burnell
Robert Busa-Fekete
Mu Cai
Glenn Cameron
Charlotte Caucheteux
Garima Chadha
Jetha Chan
Aditya Chawla
Blake Jianhang Chen
Jesse Chen
Lin Chen
Xu Chen
Derek Cheng
Tzu-hsiang Chien
Nikolai Chinaev
Yi Chou
Zhaohui Chu
Benjamin Coleman
Pooja Consul
Sam Conway-Rahman
Scott Crowell
Dylan Cutler
Vivek Dani
Samira Daruki
Anil Das
Daniel Deutsch
Nishanth Dikkala
Li Ding
Qiuhan Ding
Shenil Dodhia
Konstantin Donhauser
Tulsee Doshi
Anca Dragan
Alex Druinsky
Sahil Dua
Zoltan Egyed
Danielle Eisenbud
Daniel Eppens
Cindy Fan
Bahare Fatemi
Yassir Fathullah
Vlad Feinberg
Milen Ferev
Takumi Fujimoto
Isaac Galatzer-Levy
João Gante
Simon Geisler
Soham Ghosal
Antonious M. Girgis
Alec Go
Alhaad Gokhale
Alex Grills
Yiming Gu
Pramod Gupta
Guru Guruganesh
Raia Hadsell
Hamza Harkous
Jitendra Harlalka
Demis Hassabis
Anja Hauth
Joe Heyward
Arian Hosseini
Chih-Yang Hsia
I-Hung Hsu
Xiaopeng Huang
Yangsibo Huang
Kevin Hui
Adrian Hutter
Te I
Fotis Iliopoulos
Advait Jain
Ganesh Jawahar
Ziwei Ji
Qilin Jin
Melvin Johnson
Kandarp Joshi
Arun Kandoor
Wang-Cheng Kang
Koray Kavukcuoglu
Mehran Kazemi
Kathleen Kenealy
Amr Khalifa
Phoebe Kirk
Suraj Kothawade
Vitaly Kovalev
Neel Kovelamudi
Adam Kraft
Ravin Kumar
Harish Kuppam
Justin Lannin
Chen-Yu Lee
Seungji Lee
Dmitry Lepikhin
Dongdong Li
Qiujia Li
Valentin Liévin
Ethan Lin
Ziqian Lin
Casper Liu
Tianlin Liu
Tianqi Liu
Xin Liu
Mayank Lunayach
Min Ma
Gagan Madan
Andrii Maksai
Eric Malmi
Michal Matuszak
Daniel McDuff
Gaurav Menghani
Daniil Mirylenka
Karolis Misiunas
Vedant Misra
Andreea Mitran
Kareem Mohamed
Maksim Mukha
Eric Noland
James O’Donnell
Kate Olszewska
Bernett Orlando
Wanqiong Pan
Rina Panigrahy
Unnati Parekh
Chunjong Park
Eric Paskie
Liqian Peng
Bryce Petrini
Slav Petrov
Jonas Pfeiffer
Bilal Piot
Martyna Plomecka
Siim Poder
Octavio Ponce
Arijit Pramanik
David Racz
Anish Rajan
Michelle Ramanovich
Anand Rao
Marvin Ritter
Vitor Rodrigues
Evan Rosen
Mikołaj Rybiński
Noveen Sachdeva
Michaël E. Sander
Rohit Sathyanarayana
Sagar Savla
Samuel Schmidgall
Tal Schuster
Benoit Seguin
Andrew Sellergren
Aliaksei Severyn
Izhak Shafran
Dhruv Shah
Yuan Shangguan
Ashish Shenoy
Pradeep Shenoy
Rakesh Shivanna
Pauline Sho
Lucas Spangher
Wojciech Stokowiec
Tim Strother
Yao Su
Yinghao Sun
Mukund Sundararajan
Andrea Tacchetti
Mor Hazan Taege
Pouya Tafti
Chetan Tekur
Rahul Thapa
Madeleine Traverse
Lenart Treven
Tao Tu
Chien Te Tung
Petar Veličković
Malini Pooni Venkat
Sagar Gubbi Venkatesh
Vidya Venkiteswaran
Francesco Visin
Alex Vitvitskyi
Kiran Vodrahalli
Weiyi Wang
Xin Wang
Tris Warkentin
Jan Wassenberg
John Wieting
Lechao Xiao
Hao Xu
Yuhui Xu
Fuzhao Xue
Arun Yadav
Jun Yan
Antoine Yang
Lin Yang
Ming-Hsuan Yang
Ziyu Ying
Jae Hyeon Yoo
Sajjad Zafar
Fred Zhang
Jiageng Zhang
Jianyi Zhang
Xiaofan Zhang
Chao Zhao
David Zhou
Chen Zou
Appendix附录
Conversation format. We give an example of a conversation including thinking, function definition and function calling in Table 11.对话格式。我们在表 11 中给出了一个包含思考、函数定义和函数调用的对话示例。
Vision. We detail the vision encoder architecture in Table 10. We then illustrate how images are resized before being fed to the vision encoder in Figure 2, and detail the resizing algorithm in Algorithm 1. We display the vision benchmark scores of Gemma 4 models at low resolution () in Table 12.视觉。我们在表 10 中详细介绍了视觉编码器架构。接着,我们在图 2 中说明了图像在输入视觉编码器之前如何进行缩放,并在算法 1 中详细介绍了缩放算法。我们在表 12 中展示了 Gemma 4 模型在低分辨率 (Nmax=280N_{max}=280) 下的视觉基准测试得分。
| Total Params | ||||
| 550M | 1152 | 4304 | 16 | 27 |
| 150M | 768 | 3072 | 12 | 16 |
| Context | Formatting |
| Thinking toggle | <|think|> |
| Function declaration | <|tool>declaration:...<tool|> |
| Function call | <|tool_call>call:...<tool_call|> |
| Thinking trace | <|channel>thought …<channel|> |
| System turn | <|turn>system |
| User turn | <|turn>user |
| Model turn | <|turn>model |
| End of turn | <turn|> |
| Example of discussion: | |
| Toggle thinking mode. Declare function. User: I want you to book a train ticket for me. Model: <…> Where would you like to go? User: To Rome. Model: <…> Looking for available tickets: <function call> | |
| Model input: | |
| [BOS] <|turn>system <|think|> <|tool>declaration:search_train{…}<tool|><turn|> <|turn>user I want you to book a train ticket for me.<turn|> <|turn>model <|channel>thought …<channel|>Where would you like to go?<turn|> <|turn>user To Rome.<turn|> <|turn>model | |
| Model output: | |
| <|channel>thought …<channel|>Looking for available tickets: <|tool_call>call:search_train{from:<|"|>Athens<|"|>,to:<|"|>Rome<|"|>} <tool_call|><turn|> | |
| Gemma 4 | |||||
| 31B | 26B-A4B | 12B | E4B | E2B | |
| MMMU Pro | 75.8 | 73.2 | 67.7 | 51.4 | 43.2 |
| MATH-Vision | 83.4 | 80.3 | 76.7 | 59.2 | 53.0 |
| MedXPertQA MM | 60.7 | 55.7 | 47.4 | 28.7 | 22.5 |
| InfographicVQA | 82.8 | 77.8 | 58.7 | 54.8 | 44.6 |
| OmniDocBench 1.5 | 0.201 | 0.269 | 0.408 | 0.307 | 0.496 |