DSpark:基于置信度调度的推测解码与半自回归生成

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

深度求索 DeepSeek-AI · DeepSeek · 2026-06-27 · GitHub · DeepSpec ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

推测解码通过将草稿生成与目标验证解耦来加速大语言模型推理。虽然最近的并行草稿器能在单次前向传播中高效生成长令牌序列,但由于缺乏令牌间依赖关系,其接受率会迅速下降。此外,不加区分地验证这些扩展块会浪费关键批处理容量在高拒绝风险的令牌上,严重降低高并发服务系统的吞吐量。我们提出了 DSpark,一个统一高吞吐并行生成与自适应、负载感知验证的推测解码框架。为保持草稿质量,DSpark 采用半自回归架构——将并行主干与轻量级顺序模块耦合——引入块内依赖建模并缓解后缀衰减。为优化系统效率,DSpark 采用置信度调度验证,基于估计的前缀存活概率和引擎特定的吞吐量配置文件,动态调整每个请求的验证长度。在跨多个领域的离线基准测试中,DSpark 在接受的令牌长度上显著优于最先进的自回归和并行草稿器。当部署在 DeepSeek-V4 服务系统中并面对实时用户流量时,DSpark 成功缓解了验证浪费。与已建立的生产基线(MTP-1)相比,DSpark 在匹配吞吐量水平下将每用户生成速度提升了 60%-85%。更重要的是,通过防止在严格交互性约束下吞吐量严重下降,它实现了以前无法达到的性能层级,改变了服务系统的帕累托前沿。为促进社区进步,我们开源了 DSpark 检查点以及 DeepSpec,一个用于推测解码的算法驱动训练仓库。

Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture—coupling a parallel backbone with a lightweight sequential module—to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60%–85% at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system. To facilitate community progress, we open-source the DSpark checkpoints alongside DeepSpec, an algorithm-driven training repository for speculative decoding.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 7)

全文 · Full text(逐段中英对照)

引言 Introduction

大型语言模型(LLM)以自回归方式生成文本:每个新词元都需要基于所有先前词元进行一次完整的前向传播,这使得推理延迟与输出长度成正比。由此导致的低 GPU 利用率和用户感知的高等待时间构成了*同等贡献。

Large Language Models (LLMs) generate text autoregressively: each new token requires a full forward pass conditioned on all preceding tokens, making inference latency proportional to the output length. The resulting low GPU utilization and high user-perceived waiting time constitute * Equal contribution.

生产级 LLM 服务的主要瓶颈,尤其是在延迟敏感的场景中,例如实时对话助手和多轮智能体工作流。推测解码(Chen 等人,2023;Leviathan 等人,2023)提供了一种原则性解决方案:轻量级草稿模型提出一组候选词元,全尺寸目标模型通过拒绝采样在一次前向传播中验证整个块,接受与目标分布一致的最长前缀并附加一个奖励词元。由于验证是并行的,且接受规则精确保持目标分布,推测解码在无质量损失的情况下加速生成。草稿模型的设计决定了草稿延迟与接受率之间的权衡。早期的草稿器是自回归的(Cheng 等人,2024;Li 等人,2024b),每个位置基于先前采样的词元进行条件生成。然而,其草稿延迟随块大小线性增长,迫使这些方法使用短块和浅层架构。为了打破这一顺序瓶颈,并行草稿器(Cai 等人,2024;Chen 等人,2026;Liu 等人,2026a)作为一种有吸引力的替代方案出现:所有草稿位置在一次前向传播中生成,使得草稿延迟几乎与块大小无关。这种结构优势理论上允许并行草稿器高效生成更长的草稿块。然而,充分释放大并行草稿块的潜力引入了两个关键瓶颈——一个在生成质量方面,另一个在系统效率方面。首先,由于并行草稿器独立预测每个位置,它们无法建模块内词元间的依赖关系。这种独立性导致多模态冲突和后续位置的接受率快速衰减(Gu 等人,2018;Huang 等人,2022b)。其次,确定最优验证长度仍然是一个挑战。虽然并行生成容易产生长草稿块,但不加区分地验证所有提议词元会降低系统吞吐量,特别是在高并发工作负载下(Hu 等人,2026;Liu 等人,2024c)。理想的验证长度沿两个轴变化。在数据方面,像代码这样的结构化请求自然比开放式聊天具有更高的接受率(Abramovich 等人,2026;Xia 等人,2024)。在系统方面,在轻负载下验证额外词元几乎免费。然而,在重负载下,验证具有高拒绝风险的词元会占用关键的批处理容量,而这些容量本可以用于服务其他活跃请求(Liu 等人,2024b;Wu 等人,2025)。为了解决这些瓶颈,我们引入了 DSpark,一个将高吞吐量并行生成与自适应、负载感知验证相结合的推测解码框架。其核心是,DSpark 旨在通过两种互补机制解决草稿生成和验证中的固有权衡。• 首先,为了克服缺乏词元间依赖关系的问题,DSpark 采用半自回归架构。它保持计算昂贵的草稿骨干完全并行,仅附加一个轻量级顺序输出头来注入局部转移信息。这种设计保留了并行模型的草稿速度,同时显著减轻了后缀衰减。• 其次,为了解决系统级瓶颈,DSpark 采用置信度调度验证。通过将置信度头(估计每个位置前缀存活概率)与硬件感知调度器相结合,DSpark 动态地为每个请求定制验证长度。该调度器利用实时引擎吞吐量配置文件,仅将目标验证预算分配给具有最高预期回报的词元。我们在受控离线基准测试和生产级在线部署上广泛评估了 DSpark。在受控离线基准测试上——涵盖数学推理、代码生成和日常聊天——DSpark 始终优于强基线。具体来说,

a primary bottleneck in production LLM serving, particularly for latency-sensitive scenarios such as real-time conversational assistants and multi-turn agentic workflows. Speculative decoding (Chen et al., 2023; Leviathan et al., 2023) offers a principled solution: a lightweight draft model proposes a block of candidate tokens, and the full-size target model verifies the entire block in a single forward pass via rejection sampling, accepting the longest prefix consistent with the target distribution and appending one bonus token. Because verification is parallel and the acceptance rule preserves the target distribution exactly, speculative decoding accelerates generation without any quality loss. The design of the draft model governs the trade-off between drafting latency and acceptance rate. Early drafters are autoregressive (Cheng et al., 2024; Li et al., 2024b), conditioning each position on previously sampled tokens. However, their drafting latency grows linearly with the block size, forcing these methods to use short blocks and shallow architectures. To break this sequential bottleneck, parallel drafters (Cai et al., 2024; Chen et al., 2026; Liu et al., 2026a) have emerged as a compelling alternative: all draft positions are produced in a single forward pass, making drafting latency nearly independent of block size. This structural advantage theoretically allows parallel drafters to efficiently generate substantially longer draft blocks. However, fully unlocking the potential of large parallel draft blocks introduces two critical bottlenecks—one in generation quality, and the other in system efficiency. First, because parallel drafters predict each position independently, they cannot model inter-token dependencies within a block. This independence leads to multi-modal collisions and rapid acceptance decay at later positions (Gu et al., 2018; Huang et al., 2022b). Second, determining the optimal verification length remains a challenge. While parallel generation easily produces long draft blocks, indiscriminately verifying all proposed tokens degrades system throughput, particularly under high-concurrency workloads (Hu et al., 2026; Liu et al., 2024c). The ideal verification length varies along two axes. On the data side, structured requests like code naturally sustain higher acceptance rates than open-ended chat (Abramovich et al., 2026; Xia et al., 2024). On the system side, verifying extra tokens is nearly free under light loads. Under heavy loads, however, verifying tokens with a high rejection risk occupies critical batch capacity that could otherwise serve other active requests (Liu et al., 2024b; Wu et al., 2025). To address these bottlenecks, we introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. At its core, DSpark is designed to resolve the inherent trade-offs in draft generation and verification through two complementary mechanisms. • First, to overcome the lack of inter-token dependencies, DSpark adopts a semi-autoregressive architecture. It keeps the computationally expensive draft backbone fully parallel, appending only a lightweight serial output head to inject local transition information. This design preserves the drafting speed of parallel models while significantly mitigating suffix decay. • Second, to resolve the system-level bottleneck, DSpark employs confidence-scheduled verification. By coupling a confidence head—which estimates per-position prefix survival probabilities—with a hardware-aware scheduler, DSpark dynamically tailors the verification length for each request. This scheduler leverages real-time engine throughput profiles to route target verification budget only toward tokens with the highest expected return. We extensively evaluate DSpark across both controlled offline benchmarks and productionscale online deployments. On controlled offline benchmarks—spanning mathematical reasoning, code generation, and daily chat—DSpark consistently outperforms strong baselines. Specifically, 2

在 Qwen3-4B、8B 和 14B 目标模型(Yang 等人,2025)上,与自回归的 Eagle3(Li 等人,2026b)相比,宏平均接受长度分别提高了 30.9%、26.7%和 30.0%,与并行的 DFlash(Chen 等人,2026)相比,分别提高了 16.3%、18.4%和 18.3%。除了顶级指标,我们的细粒度逐位置分析揭示了不同草稿器的不同生成特征,实证展示了 DSpark 如何成功地将并行模型的高初始词元容量与自回归模型的后缀连贯性结合起来。除了离线评估,我们将 DSpark 部署在 DeepSeek-V4(DeepSeek-AI,2026)服务系统中,以评估其在实时用户流量下的性能。与之前的 MTP-1 生产基线(DeepSeek-AI,2024)相比,DSpark 显著拓宽了系统的运行范围。具体来说,在匹配的总吞吐量容量下,它始终将每用户生成速度提高 60%–85%(V4Flash)和 57%–78%(V4-Pro)。此外,在严格的服务水平协议(SLA)下,当基线的容量严重恶化时——例如 Flash 的 120 TPS 和 Pro 的 50 TPS——DSpark 减轻了验证开销以保持稳健的吞吐量。通过克服这一性能悬崖,DSpark 解锁了以前无法实现的严格交互性层级,有效地移动了 LLM 服务的帕累托前沿。为了促进开源社区的集体进步,我们公开提供我们的工件。具体来说,我们发布了针对 DeepSeek-V4-Flash(预览版)和 DeepSeek-V4-Pro(预览版)模型训练的 DSpark 检查点。此外,我们开源了 DeepSpec,一个算法驱动的训练仓库,包括 Eagle3、DFlash 和 DSpark。这些工件旨在支持关于高效 LLM 服务的进一步研究。

across the Qwen3-4B, 8B, and 14B target models (Yang et al., 2025), it improves the macro-average accepted length over the autoregressive Eagle3 (Li et al., 2026b) by 30.9%, 26.7%, and 30.0%, and over the parallel DFlash (Chen et al., 2026) by 16.3%, 18.4%, and 18.3%, respectively. Beyond topline metrics, our fine-grained position-wise analysis reveals the distinct generation characteristics of different drafters, empirically demonstrating how DSpark successfully combines the high initial-token capacity of parallel models with the suffix coherence of autoregressive models. Beyond offline evaluation, we deployed DSpark within the DeepSeek-V4 (DeepSeek-AI, 2026) serving system to assess its performance under live user traffic. Compared to the prior MTP-1 production baseline (DeepSeek-AI, 2024), DSpark significantly broadens the system’s operational envelope. Specifically, it consistently accelerates per-user generation speeds by 60%–85% (V4Flash) and 57%–78% (V4-Pro) at matched aggregate throughput capacities. Furthermore, under strict Service Level Agreements (SLAs) where the baseline’s capacity deteriorates severely—such as 120 TPS for Flash and 50 TPS for Pro—DSpark mitigates verification overhead to maintain robust throughput. By overcoming this performance cliff, DSpark unlocks strict interactivity tiers that were previously unattainable, effectively shifting the Pareto frontier of LLM serving. To foster collective advancement within the open-source community, we are making our artifacts publicly available. Specifically, we release the trained DSpark checkpoints for both the DeepSeek-V4-Flash (preview) and DeepSeek-V4-Pro (preview) models. Furthermore, we open-source DeepSpec, an algorithm-driven training repository, including Eagle3, DFlash and DSpark. These artifacts are intended to support further research on efficient LLM serving.

背景 Background

2.1. 推测解码 自回归语言模型每次前向传播生成一个词元,使得推理延迟与输出长度成正比。推测解码(Chen 等人,2023;Ge 等人,2022;Leviathan 等人,2023)使用轻量级草稿模型𝑀𝑑加速目标模型𝑀𝑡的推理。在每个解码周期,草稿模型提出𝛾个候选词元𝑥1, . . . , 𝑥𝛾。目标模型在一次前向传播中验证所有候选词元,接受与其自身分布一致的最长前缀。具体来说,在每个草稿位置𝑘,目标模型计算其自身分布𝑝𝑡𝑘,并将其与草稿分布𝑝𝑘𝑑进行比较。词元𝑥𝑘以概率 min(1, 𝑝𝑡𝑘(𝑥𝑘)/𝑝𝑘𝑑(𝑥𝑘))被接受。验证从左到右进行:在位置𝑘的第一次拒绝会丢弃所有后续词元𝑥𝑘+1, . . . , 𝑥𝛾,无论其质量如何。令𝜏表示每个周期接受的词元数,𝑇draft 和𝑇verify 分别表示草稿和验证前向传播的挂钟时间。每个生成词元的平均延迟为:𝐿 = (𝑇draft + 𝑇verify) / 𝜏。

2.1. Speculative Decoding Autoregressive language models generate one token per forward pass, making inference latency proportional to output length. Speculative decoding (Chen et al., 2023; Ge et al., 2022; Leviathan et al., 2023) accelerates the inference of a target model 𝑀𝑡 using a lightweight draft model 𝑀𝑑 . At each decoding cycle, the draft model proposes 𝛾 candidate tokens 𝑥1 , . . . , 𝑥 𝛾 . The target model verifies all candidates in a single forward pass, accepting the longest prefix consistent with its own distribution. Concretely, at each draft position 𝑘, the target model computes its own distribution 𝑝𝑡𝑘 and compares it against the draft distribution 𝑝𝑘𝑑 . The token 𝑥 𝑘 is accepted with probability min(1, 𝑝𝑡𝑘 ( 𝑥 𝑘 )/ 𝑝𝑘𝑑 ( 𝑥 𝑘 )). Verification proceeds left to right: the first rejection at position 𝑘 discards all subsequent tokens 𝑥 𝑘+1 , . . . , 𝑥 𝛾 , regardless of their quality. Let 𝜏 denote the number of accepted tokens per cycle, and let 𝑇draft and 𝑇verify be the wall-clock times of the drafting and verification passes, respectively. The average latency per generated token is: 𝑇draft + 𝑇verify 𝐿= . (1) 𝜏

因此,提高加速比归结为三个杠杆:降低𝑇draft(更快地草稿),提高𝜏(更好地草稿),或降低有效的𝑇verify(更智能地验证)。2.2. 草稿器架构 草稿模型的设计决定了𝑇draft 和𝜏之间的权衡。现有方法分为两类。

Improving speedup therefore reduces to three levers: lowering 𝑇draft (draft faster), raising 𝜏 (draft better), or reducing the effective 𝑇verify (verify smarter). 2.2. Drafter Architectures The design of the draft model determines how 𝑇draft and 𝜏 trade off. Existing approaches fall into two categories. 3

自回归草稿器。自回归草稿器顺序生成草稿词元,每个位置基于先前采样的词元进行条件生成(DeepSeek-AI,2024;Li 等人,2024b,c, 2026b;Zhang 等人,2025)。这种显式依赖提供了强大的建模能力,但草稿成本随块大小线性增长:𝑇draft ∝ 𝛾,这迫使自回归草稿器使用小的𝛾和浅层架构以保持𝑇draft 较低。为了补偿短块,基于树的验证(Miao 等人,2024)将候选词元扩展成树,并通过树注意力验证多条路径,但大量验证词元降低了整体服务吞吐量。并行草稿器。并行草稿器在一次前向传播中生成所有𝛾个草稿词元,使得𝑇draft 几乎与块大小无关(Cai 等人,2024;Chen 等人,2026;Li 等人,2025a;Liu 等人,2026a;Sandler 等人,2026)。这允许使用更大的块(例如,𝛾=16)而不会成比例地增加延迟。其中,DFlash(Chen 等人,2026)是一种最先进的并行草稿器,它通过从目标模型提取的丰富上下文特征(KV 注入)来条件化其草稿模型。在预填充期间,来自一组目标层{𝑙1, . . . , 𝑙𝑚}的隐藏状态被拼接并投影到草稿隐藏空间:𝐻ctx = RMSNorm(𝑊𝑐[𝐻(𝑙1); . . . ; 𝐻(𝑙𝑚)]),其中𝑊𝑐 ∈ R𝑑×𝑚𝑑是一个共享投影。这些上下文特征通过沿键和值的序列维度与草稿块表示拼接,注入到每个草稿层:𝐾𝑖 = [𝑊𝑖𝐾𝐻ctx; 𝑊𝑖𝐾𝐻𝑑], 𝑉𝑖 = [𝑊𝑖𝑉𝐻ctx; 𝑊𝑖𝑉𝐻𝑑]。块内的所有位置彼此双向关注,并关注注入的目标上下文。草稿模型共享目标模型的嵌入层和语言建模头(均冻结)。它接受一个锚定词元 1 后跟𝛾个掩码词元嵌入作为输入,并在一次前向传播中为所有掩码位置生成 logits。由于无论块大小如何,草稿只需要一次前向传播,因此在相同的延迟预算下,DFlash 可以比自回归草稿器采用更深的架构和更大的块。

Autoregressive drafters. Autoregressive drafters generate draft tokens sequentially, conditioning each position on previously sampled tokens (DeepSeek-AI, 2024; Li et al., 2024b,c, 2026b; Zhang et al., 2025). This explicit dependency gives strong modeling capacity, but the drafting cost grows linearly with block size: 𝑇draft ∝ 𝛾 , which forces autoregressive drafters to use small 𝛾 and shallow architectures to keep 𝑇draft low. To compensate for the short block, tree-based verification (Miao et al., 2024) expands candidates into a tree and verifies multiple paths via tree attention, but the large number of verification tokens reduces overall serving throughput. Parallel drafters. Parallel drafters produce all 𝛾 draft tokens in a single forward pass, making 𝑇draft nearly independent of the block size (Cai et al., 2024; Chen et al., 2026; Li et al., 2025a; Liu et al., 2026a; Sandler et al., 2026). This allows substantially larger blocks (e.g., 𝛾 =16) without proportionally increasing latency. Among them, DFlash (Chen et al., 2026) is a state-of-the-art parallel drafter, which conditions its draft model on rich context features extracted from the target model (KV injection). During prefill, hidden states from a set of target layers { 𝑙1 , . . . , 𝑙 𝑚 } are concatenated and projected into the draft hidden space:  𝐻ctx = RMSNorm 𝑊𝑐 [ 𝐻 ( 𝑙1 ) ; . . . ; 𝐻 ( 𝑙𝑚 ) ] , (2) where 𝑊𝑐 ∈ R𝑑 × 𝑚𝑑 is a shared projection. These context features are injected into every draft layer by concatenating them with the draft block representations along the sequence dimension of keys and values: 𝐾𝑖 = [𝑊𝑖𝐾 𝐻ctx ; 𝑊𝑖𝐾 𝐻𝑑 ], 𝑉𝑖 = [𝑊𝑖𝑉 𝐻ctx ; 𝑊𝑖𝑉 𝐻𝑑 ]. (3) All positions within a block attend bidirectionally to each other and to the injected target context. The draft model shares the target model’s embedding layer and language modeling head (both frozen). It takes as input the embedding of an anchor token1 followed by 𝛾 mask token embeddings, and produces logits for all mask positions in a single forward pass. Since drafting requires only a single forward pass regardless of block size, DFlash can afford deeper architectures and larger blocks than autoregressive drafters under the same latency budget.

架构 Architecture

DSpark 的概述如图 1 所示。回顾公式 1,推测解码的每词元延迟为𝐿 = (𝑇draft + 𝑇verify)/𝜏。自回归草稿器实现了高𝜏,但付出了𝑇draft ∝ 𝛾的代价;并行草稿器将𝑇draft 压缩为单次前向传播,但牺牲了𝜏,因为每个位置是独立预测的。同时,固定长度验证将𝑇verify 浪费在几乎肯定被拒绝的低置信度后缀词元上。DSpark 通过两个互补组件解决了这些局限性:• 半自回归生成(第 3.1 节)。并行骨干处理草稿计算的主体,使𝑇draft 几乎与𝛾无关。轻量级顺序块随后在草稿词元之间注入依赖关系,以最小的额外延迟提高𝜏。• 置信度调度验证(第 3.2 节)。置信度头估计每个位置的接受概率,硬件感知调度器使用这些估计来修剪低置信度后缀词元,减少不必要的验证计算。

The overview of DSpark is shown in Figure 1. Recall from Equation 1 that the per-token latency of speculative decoding is 𝐿 = (𝑇draft + 𝑇verify )/𝜏. Autoregressive drafters achieve high 𝜏 but pay 𝑇draft ∝ 𝛾 ; parallel drafters collapse 𝑇draft to a single pass but sacrifice 𝜏 because each position is predicted independently. Meanwhile, fixed-length verification wastes 𝑇verify on low-confidence suffix tokens that are almost certain to be rejected. DSpark addresses these limitations with two complementary components: • Semi-autoregressive generation (Section 3.1). A parallel backbone handles the bulk of draft computation, which keeps 𝑇draft nearly independent of 𝛾 . A lightweight sequential block then injects dependency among draft tokens, improving 𝜏 at minimal additional latency. • Confidence-scheduled verification (Section 3.2). A confidence head estimates per-position acceptance probabilities, and a hardware-aware scheduler uses these estimates to prune low-confidence suffix tokens, cutting unnecessary verification compute. 1We use the terms anchor token and bonus token interchangeably in this paper to denote the final token generated

图 1 | DSpark 架构和解码周期。给定提示词元 ABC,目标模型执行一步生成下一个词元 D,作为草稿阶段的锚点。以 D 为输入,DSpark 使用重型并行骨干和轻量级顺序头生成草稿词元 EFGH 及其对应的置信度分数𝑐1–𝑐4。硬件感知前缀调度器随后评估这些分数以保留前缀 EFG 并丢弃低置信度词元 H。最后,目标模型并行验证调度后的前缀。如图所示,E 和 F 被接受,而 G 被拒绝,促使模型生成修正词元 G∗以完成当前轮次。结合这两个组件,DSpark 实现了更好的草稿和更智能的验证。我们下面详细说明每个组件。3.1. 半自回归生成 并行草稿器在一次前向传播中生成所有𝛾个草稿 logits,因此每个预测不能基于块内其他位置采样的词元进行条件化。当上下文允许多个合理的延续时,例如“of course”和“no problem”,并行草稿器可能产生不连贯的组合,如“of problem”或“no course”,因为每个位置对所有可能的前驱进行边缘化,而不是基于实际采样的那个进行条件化(Gu 等人,2018;Huang 等人,2022a)。因此,接受率沿块快速衰减,浪费了草稿和验证计算。因此,我们采用半自回归结构,将草稿生成分为两个阶段:并行阶段。并行骨干(在我们的实例中为 DFlash(Chen 等人,2026))对整个块运行一次前向传播,产生隐藏状态ℎ1, . . . , ℎ𝛾和基础 logits 𝑈1, . . . , 𝑈𝛾。我们对原始 DFlash 骨干进行了微小修改:不是输入一个锚定词元加上𝛾个掩码词元并仅预测掩码位置,而是

Figure 1 | The DSpark architecture and decoding cycle. Given prompt tokens ABC , the target model executes one step to generate the next token D , which serves as the anchor for the drafting phase. Using D as the input, DSpark employs a heavy parallel backbone and a lightweight sequential head to generate draft tokens EFGH along with their corresponding confidence scores 𝑐1 –𝑐4 . The Hardware-Aware Prefix Scheduler then evaluates these scores to retain the prefix EFG and drop the low-confidence token H . Finally, the target model verifies the scheduled prefix in parallel. As illustrated, E and F are accepted while G is rejected, prompting the model to generate a corrected token G∗ to complete the current round. In combination, the two components let DSpark draft better and verify smarter. We detail each below. 3.1. Semi-Autoregressive Generation A parallel drafter produces all 𝛾 draft logits in one forward pass, so each prediction cannot condition on tokens sampled elsewhere in the block. When the context admits multiple plausible continuations, e.g., “of course” and “no problem”, a parallel drafter may produce incoherent combinations such as “of problem” or “no course”, because each position marginalizes over all possible predecessors rather than conditioning on the one actually sampled (Gu et al., 2018; Huang et al., 2022a). Acceptance rate thus decays rapidly along the block, wasting both draft and verification compute. We therefore adopt a semi-autoregressive structure that splits draft generation into two stages: Parallel stage. A parallel backbone (in our instantiation, DFlash (Chen et al., 2026)) runs a single forward pass over the entire block, producing hidden states ℎ1 , . . . , ℎ𝛾 and base logits 𝑈1 , . . . , 𝑈𝛾 . We make only a minor modification to the original DFlash backbone: instead of feeding an anchor token plus 𝛾 mask tokens and predicting only the mask positions, we treat 5

将锚定词元本身视为第一个预测位置,因此𝛾个输入词元(锚定词元 + 𝛾-1 个掩码)产生𝛾个草稿 logits。这减少了草稿计算,同时保持了相似的草稿质量。顺序阶段。顺序阶段为基础 logits 补充了一个依赖于前缀的转移偏置𝐵𝑘(𝑥0, 𝑥<𝑘, 𝑥𝑘),允许每个草稿位置基于块内先前采样的词元进行条件化。顺序阶段不是定义一个全局归一化的能量模型,而是通过自回归分解引入因果块分布:𝑃(𝑋|𝑥0) = exp(𝑈𝑘(𝑣) + 𝐵𝑘(𝑥0, 𝑥<𝑘, 𝑣)) / Σ𝑢∈V exp(𝑈𝑘(𝑢) + 𝐵𝑘(𝑥0, 𝑥<𝑘, 𝑢))。这里,𝑥0 表示来自上一个验证周期的锚定词元,𝑈𝑘是并行骨干在位置𝑘产生的基础 logit 向量,V 是词汇表。在推理时,顺序块根据𝑝𝑘(·|𝑥0, 𝑥<𝑘)从左到右采样。由于此采样过程本质上是顺序的,块必须在计算上轻量(𝑇sequential ≪ 𝑇parallel),以便整体草稿延迟仍由并行阶段主导。我们在下面描述顺序块的两种实例化。• 马尔可夫头。最简单的实例化将𝐵𝑘限制为仅依赖于紧邻的前一个词元,将其简化为一个一阶转移𝐵(𝑥𝑘−1, 𝑥𝑘)。原则上这是一个完整的𝑉×𝑉矩阵𝐵;我们通过低秩分解𝐵 = 𝑊1𝑊2 来近似它,其中𝑊1 ∈ R𝑉×𝑟且𝑊2 ∈ R𝑟×𝑉。给定前一个词元𝑥𝑘−1,位置𝑘的转移偏置为:𝐵(𝑥𝑘−1, ·) = 𝑊1[𝑥𝑘−1]𝑊2 ∈ R𝑉,其中𝑊1 作为嵌入查找表,𝑊2 作为 logit 投影。低秩分解(默认𝑟=256)保持了存储和每步计算的小规模,使得顺序循环即使对于大词汇表也高效。回到前面的例子:一旦位置 1 采样了“of”,马尔可夫头在位置 2 提升“course”并抑制“problem”,从而减轻了跨模态冲突。• RNN 头。马尔可夫头在一步之外是无记忆的——位置𝑘无法访问𝑥𝑘−1 之前的词元。RNN 头通过维护一个循环状态𝑠𝑘来放松这一限制,该状态累积块内的完整前缀历史。在每一步,模块将当前状态𝑠𝑘−1 ∈ R𝑟、前一个词元嵌入𝑊1[𝑥𝑘−1] ∈ R𝑟和骨干隐藏状态ℎ𝑘 ∈ R𝑑拼接成输入向量𝑧𝑘 = [𝑠𝑘−1; 𝑊1[𝑥𝑘−1]; ℎ𝑘] ∈ R2𝑟+𝑑,然后应用一个门控更新:𝑠𝑘 = 𝜎(𝑊𝑔𝑧𝑘) ⊙ 𝑠𝑘−1 + (1 − 𝜎(𝑊𝑔𝑧𝑘)) ⊙ tanh(𝑊𝑐𝑧𝑘), 𝐵𝑘(𝑥<𝑘, ·) = 𝑊2⊤ tanh(𝑊𝑜𝑧𝑘),其中𝑊𝑔, 𝑊𝑐, 𝑊𝑜 ∈ R(2𝑟+𝑑)×𝑟通过一个线性投影联合参数化,该投影被分割为门、候选和输出组件。状态𝑠0 初始化为零。3.2. 置信度调度验证 半自回归架构使 DSpark 能够高效生成大的草稿块。然而,产生更多的草稿词元并不自动转化为更高的端到端加速比。不加区分地验证整个草稿块实际上会降低整体系统吞吐量,特别是在高并发场景中(Hu 等人,2026;Liu 等人,2024c)。

the anchor itself as the first prediction position, so 𝛾 input tokens (anchor + 𝛾 −1 masks) yield 𝛾 draft logits. This reduces draft computation while maintaining similar draft quality. Sequential stage. The sequential stage supplements the base logits with a prefix-dependent transition bias 𝐵𝑘 ( 𝑥0 , 𝑥 <𝑘 , 𝑥 𝑘 ), allowing each draft position to condition on previously sampled tokens within the block. Rather than defining a globally normalized energy model, the sequential stage induces a causal block distribution through an autoregressive factorization: 𝑃 ( 𝑋 | 𝑥0 ) =

这种性能瓶颈源于两个相互作用的因素。首先,在数据方面,草稿接受率在不同领域固有地变化:像代码这样的结构化文本自然产生高接受率,而开放式聊天具有显著较低的接受率(Abramovich 等人,2026;Xia 等人,2024)。其次,在系统方面,验证额外词元的实际成本严格取决于引擎负载。在系统轻负载下,即使被拒绝,额外验证也带来最小惩罚。然而,在高并发部署中,每次不必要的验证都会占用目标模型的批处理容量,而这些容量本可以用于服务其他活跃请求(Liu 等人,2024b;Wu 等人,2025)。因此,充分释放大草稿块的潜力需要一个统一的机制,将目标模型计算仅路由到具有正预期回报的词元。DSpark 通过将置信度头(第 3.2.1 节)与硬件感知前缀调度器(第 3.2.2 节)相结合来实现这一点,置信度头预测前缀存活概率,调度器基于当前系统负载动态确定最优验证长度。3.2.1. 置信度头 受 Huang 等人(2024);Wang 等人(2026)的启发,置信度头为每个草稿位置𝑘输出一个标量估计𝑐𝑘 ∈ (0, 1)。关键的是,𝑐𝑘建模了草稿词元在位置𝑘将在目标验证中存活的概率,前提是块中所有前面的词元已被接受。该架构采用轻量级线性投影后跟 sigmoid 函数:𝑐𝑘 = 𝜎(𝑤⊤[ℎ𝑘; 𝑊1[𝑥𝑘−1]]),其中ℎ𝑘是骨干的隐藏状态,𝑊1[𝑥𝑘−1]是来自前一个草稿词元的马尔可夫嵌入。我们使用每步分析接受率𝑐∗𝑘来监督𝑐𝑘。该率由草稿分布𝑝𝑘𝑑和目标分布𝑝𝑡𝑘之间的总变差距离决定:𝑐∗𝑘 = 1 − 1/2 ∥𝑝𝑘𝑑 − 𝑝𝑡𝑘∥1。事后校准。与基于阈值的验证启发式方法(Huang 等人,2024;Li 等人,2024b;Zhang 等人,2026b)不同,这些方法仅需要置信度分数正确地对草稿词元质量进行排序,我们的硬件感知调度方法(在第 3.2.2 节中详述)精确地需要累积接受概率的绝对值来计算预期接受长度𝜏。由于神经置信度估计通常过于自信(Guo 等人,2017;Ovadia 等人,2019),直接使用原始置信度分数会扭曲吞吐量估计,导致次优调度。为了解决这个问题,我们引入了顺序温度缩放(STS)。由于每个𝑐𝑖建模一个条件概率,链式法则规定草稿前缀被接受的联合概率分解为累积乘积Π𝑖≤𝑘 𝑐𝑖。使用一个保留的验证集,STS 从左到右连续校准这个联合概率。具体来说,在每个位置𝑘 ∈ {1, . . . , 𝛾},我们执行一个简单的一维网格搜索,找到最小化累积乘积的期望校准误差(ECE)(Naeini 等人,2015)的最优温度标量,同时保持所有前面位置已校准的分数固定。关键的是,温度缩放是一个保序变换:它修正预测概率以匹配经验接受率,而不破坏置信度头学习的相对草稿词元排名。3.2.2. 硬件感知前缀调度器

exp(𝑈𝑘 ( 𝑣) + 𝐵𝑘 ( 𝑥0 , 𝑥 <𝑘 , 𝑣)) . 𝑢 ∈ V exp(𝑈 𝑘 ( 𝑢) + 𝐵 𝑘 ( 𝑥 0 , 𝑥 <𝑘 , 𝑢))

算法 1 硬件感知前缀调度器 要求:活跃请求𝑟 ∈ {1, . . . , 𝑅};每个请求的置信度序列𝑐𝑟,1, . . . , 𝑐𝑟,𝛾;分析得到的步曲线 SPS(𝐵) 确保:每个请求选择的前缀长度ℓ1∗, . . . , ℓ𝑅∗ 1: for 𝑟 = 1 to 𝑅 do 2: 计算前缀存活概率:𝑎𝑟,𝑗 ← Π𝑖≤𝑗 𝑐𝑟,𝑖 对于𝑗 = 1, . . . , 𝛾 3: end for 4: 构建候选空间 E ← {(𝑟, 𝑗) | 𝑎𝑟,𝑗 > 0}并按𝑎𝑟,𝑗降序排序 5: 初始化状态:ℓ𝑟 ← 0 对所有𝑟;批大小𝐵 ← 𝑅;期望接受数𝜏∗ ← 𝑅 6: 初始化跟踪:Θbest ← 𝑅 · SPS(𝑅);选择长度ℓ𝑟∗ ← 0 对所有𝑟 7: for each (𝑟, 𝑗) ∈ E in sorted order do 8: ℓ𝑟 ← 𝑗; 𝐵 ← 𝐵 + 1; 𝜏∗ ← 𝜏∗ + 𝑎𝑟,𝑗 9: 当前吞吐量Θ ← 𝜏∗ · SPS(𝐵) 10: if Θ > Θbest then 11: Θbest ← Θ; 更新选择长度ℓ𝑟∗ ← ℓ𝑟 12: else 13: break 14: end if 15: end for 16: return (ℓ1∗, . . . , ℓ𝑅∗) 达到Θbest 先前的方法(Huang 等人,2024;Li 等人,2024b)通常对置信度分数应用静态阈值来确定验证长度。虽然在孤立的单请求假设下有效,但在高并发生产系统中,静态阈值可能是次优的,因为验证草稿词元的效用严重依赖于当前系统负载。为了解决这个问题,我们将验证长度选择表述为一个全局吞吐量最大化问题(算法 1)。考虑一批𝑅个活跃请求。对于请求𝑟,令𝑐𝑟,1, . . . , 𝑐𝑟,𝛾为逐位置置信度估计,ℓ𝑟 ∈ {0, . . . , 𝛾}表示调度的验证长度。由于推测解码仅作为连续前缀动态接受草稿词元,位置𝑗处词元的存活概率是累积乘积𝑎𝑟,𝑗 = Π𝑖≤𝑗 𝑐𝑟,𝑖。在一个验证步骤中,发送给目标模型的总批大小(以词元计)为𝐵 = Σ𝑟=1^𝑅 (1 + ℓ𝑟),期望成功接受的词元数为𝜏 = Σ𝑟=1^𝑅 (1 + Σ𝑗=1^ℓ𝑟 𝑎𝑟,𝑗)。令 SPS(𝐵)表示引擎吞吐量,以每秒步数计,对于给定的前向传播批大小𝐵。关键的是,这个容量曲线在引擎初始化期间分析一次,并存储为轻量级成本表。然后,我们的调度器旨在通过动态选择验证长度ℓ1, . . . , ℓ𝑅来最大化期望的系统级词元吞吐量Θ = 𝜏 · SPS(𝐵)。尽管找到Θ的全局最大值看起来是一个组合搜索,但目标结构允许一个高效的贪心解。由于𝑎𝑟,𝑗关于𝑗单调非增(即𝑎𝑟,𝑗 ≤ 𝑎𝑟,𝑗−1),将请求𝑟的验证长度从𝑗−1 扩展到𝑗的期望接受词元边际增益正好是𝑎𝑟,𝑗。这种单调性确保全局按𝑎𝑟,𝑗对候选词元排序自然尊重块内前缀依赖关系。因此,如果总验证批大小𝐵是固定的,最优分配{ℓ𝑟}将通过从所有{𝑎𝑟,𝑗}的全局池中贪心地选择具有最高存活概率的草稿词元来确定。基于这一见解,优化可以沿着这个贪心接纳路径进行评估。

Here, 𝑥0 denotes the anchor token from the previous verification cycle, 𝑈𝑘 is the base logit vector produced by the parallel backbone at position 𝑘, and V is the vocabulary. At inference time, the sequential block samples left to right according to 𝑝𝑘 (· | 𝑥0 , 𝑥 <𝑘 ). Because this sampling process is inherently sequential, the block must be computationally lightweight (𝑇sequential ≪ 𝑇parallel ) so that the overall draft latency remains dominated by the parallel stage. We describe two instantiations of the sequential block below. • Markov head. The simplest instantiation restricts 𝐵𝑘 to depend only on the immediately preceding token, reducing it to a first-order transition 𝐵 ( 𝑥 𝑘 −1 , 𝑥 𝑘 ). In principle this is a full 𝑉 × 𝑉 matrix 𝐵; we approximate it with a low-rank factorization 𝐵 = 𝑊1𝑊2 , where 𝑊1 ∈ R𝑉 × 𝑟 and 𝑊2 ∈ R𝑟 ×𝑉 . Given the preceding token 𝑥 𝑘 −1 , the transition bias for position 𝑘 is: 𝐵 ( 𝑥 𝑘 −1 , · ) = 𝑊1 [ 𝑥 𝑘 −1 ] 𝑊2 ∈ R𝑉 , (5) where 𝑊1 serves as an embedding lookup table and 𝑊2 as a logit projection. The low-rank factorization (𝑟 =256 by default) keeps both storage and per-step compute small, making the sequential loop efficient even for large vocabularies. Returning to the earlier example: once position 1 samples “of”, the Markov head boosts “course” and suppresses “problem” at position 2, which mitigates the cross-mode collision. • RNN head. The Markov head is memoryless beyond one step—position 𝑘 cannot access tokens before 𝑥 𝑘 −1 . The RNN head relaxes this by maintaining a recurrent state 𝑠𝑘 that accumulates the full prefix history within a block. At each step, the module concatenates the current state 𝑠𝑘 −1 ∈ R𝑟 , the previous token embedding 𝑊1 [ 𝑥 𝑘 −1 ] ∈ R𝑟 , and the backbone hidden ℎ𝑘 ∈ R𝑑 into an input vector 𝑧 𝑘 = [ 𝑠𝑘 −1 ; 𝑊1 [ 𝑥 𝑘 −1 ]; ℎ𝑘 ] ∈ R2𝑟+𝑑 , then applies a single gated update:  𝑠𝑘 = 𝜎 (𝑊𝑔 𝑧 𝑘 ) ⊙ 𝑠𝑘 −1 + 1 − 𝜎 (𝑊𝑔 𝑧 𝑘 ) ⊙ tanh(𝑊𝑐 𝑧 𝑘 ), (6) 𝐵𝑘 ( 𝑥 <𝑘 , · ) = 𝑊2⊤ tanh(𝑊𝑜 𝑧 𝑘 ), where 𝑊𝑔 , 𝑊𝑐 , 𝑊𝑜 ∈ R (2𝑟+𝑑 ) ×𝑟 are jointly parameterized by a single linear projection that is split into gate, candidate, and output components. The state 𝑠0 is initialized to zero. 3.2. Confidence-Scheduled Verification The semi-autoregressive architecture enables DSpark to generate large draft blocks efficiently. However, producing more draft tokens does not automatically translate to higher end-to-end speedups. Indiscriminately verifying the full draft block can actually degrade overall system throughput, especially in high-concurrency scenarios (Hu et al., 2026; Liu et al., 2024c).

我们首先将所有有效的前缀扩展按存活概率降序全局排序。为了动态确定最优目标批大小𝐵,我们从这个排序池中逐步接纳词元,通过从预分析的成本表中进行𝑂(1)查找来更新期望吞吐量Θ。无损推测解码严格要求非预期属性:接纳决策不能依赖于未来的候选词元(Chen 等人,2023;Leviathan 等人,2023)。由于我们的置信度头依赖于先前采样词元的马尔可夫特征,计算下一个存活概率𝑎𝑟,𝑘+1 明确需要实例化的候选词元𝑥𝑟,𝑘。因此,回顾性全局搜索会无意中将𝑥𝑟,𝑘泄露到步骤𝑘的接纳决策中,引入选择偏差(我们在附录 A 中提供了一个具体反例来证明这一理论违反)。为了强制严格因果性,调度器(算法 1)采用了一种早停机制。通过在吞吐量下降时(Θ ≤ Θbest)立即中断贪心搜索,截断决策仅依赖于处理到该确切步骤的前缀。这将接纳事件与未来词元隔离开,确保精确的目标分布恢复。注意,当且仅当目标Θ是单峰的时,这种逐步早停才能产生全局最大吞吐量,这隐含地假设了平滑衰减的硬件容量曲线。我们在第 5.2 节中讨论了真实世界非平滑 SPS 特性和异步系统管道所需的工程适配。3.3. 训练 在训练期间,我们从每个目标序列中随机采样多个锚定位置,形成𝛾词元块作为训练数据。目标模型在整个训练过程中冻结;草稿模型共享其嵌入层和语言建模头并保持冻结,仅更新骨干草稿器、顺序块和置信度头。训练目标由三项组成:交叉熵损失 Lce、分布匹配损失 Ltv 和置信度损失 Lconf。所有三项都按位置加权,权重为𝑤𝑘 = exp(−(𝑘−1)/𝛾)(Chen 等人,2026),这强调了在基于前缀的验证下对期望接受长度贡献更大的早期块位置。交叉熵损失 Lce 训练草稿器预测正确的下一个词元:Lce = −Σ𝑘 𝑤𝑘 log 𝑝𝑘𝑑(𝑥𝑘∗),其中𝑥𝑘∗是真实词元,𝑝𝑘𝑑是草稿分布。分布匹配损失 Ltv 惩罚草稿和目标分布之间的总变差距离:Ltv = Σ𝑘 𝑤𝑘 · (1/2 ∥𝑝𝑘𝑑 − 𝑝𝑘𝑡∥1)。由于总变差距离是接受率的直接代理:每步接受概率等于 1 − 1/2 ∥𝑝𝑑 − 𝑝𝑡∥1(Leviathan 等人,2023),最小化 Ltv 直接最大化期望接受率。置信度损失 Lconf 是一个二元交叉熵,训练置信度头预测来自公式 8 的软接受标签𝑐∗𝑘:Lconf = −Σ𝑘 𝑤𝑘 [𝑐∗𝑘 log 𝑐𝑘 + (1 − 𝑐∗𝑘) log(1 − 𝑐𝑘)]。总体目标是三项的加权组合(默认权重𝛼ce = 0.1, 𝛼tv = 0.9, 𝛼conf = 1.0):L = 𝛼ce Lce + 𝛼tv Ltv + 𝛼conf Lconf。

This performance bottleneck stems from two interacting factors. First, on the data side, draft acceptance rates inherently vary across domains: structured text like code naturally yields high acceptance, whereas open-ended chat has significantly lower acceptance (Abramovich et al., 2026; Xia et al., 2024). Second, on the system side, the actual cost of verifying an extra token depends strictly on the engine load. Under light system load, an extra verification incurs minimal penalty even if rejected. However, under high-concurrency deployments, every unnecessary verification occupies target model batch capacity that could otherwise serve other active requests (Liu et al., 2024b; Wu et al., 2025). Therefore, fully unlocking the potential of large draft blocks requires a unified mechanism that routes target model compute only toward tokens with a positive expected return. DSpark achieves this by coupling a confidence head (Section 3.2.1) that predicts prefix survival probabilities, with a hardware-aware prefix scheduler (Section 3.2.2) that dynamically determines the optimal verification lengths based on current system load. 3.2.1. Confidence Head Drawing inspiration from Huang et al. (2024); Wang et al. (2026), the confidence head outputs a scalar estimate 𝑐𝑘 ∈ (0, 1) for each draft position 𝑘. Crucially, 𝑐𝑘 models the conditional probability that the draft token at position 𝑘 will survive target verification, given that all preceding tokens in the block have been accepted. The architecture features a lightweight linear projection followed by a sigmoid function:  𝑐𝑘 = 𝜎 𝑤⊤ [ ℎ𝑘 ; 𝑊1 [ 𝑥 𝑘 −1 ]] , (7) where ℎ𝑘 is the hidden state of the backbone and 𝑊1 [ 𝑥 𝑘 −1 ] is the Markov Embedding from the previous draft token. We supervise 𝑐𝑘 using the analytical acceptance rate per-step 𝑐∗𝑘 . This rate is determined by the total variation distance between the draft distribution 𝑝𝑘𝑑 and the target distribution 𝑝𝑡𝑘 : 𝑐∗𝑘 = 1 − 12 ∥ 𝑝𝑘𝑑 − 𝑝𝑡𝑘 ∥ 1 . (8) Post-hoc Calibration. Unlike threshold-based verification heuristics (Huang et al., 2024; Li et al., 2024b; Zhang et al., 2026b), which only require confidence scores to correctly rank draft token qualities, our hardware-aware scheduling approach (detailed in Section 3.2.2) precisely requires the absolute magnitudes of the cumulative acceptance probabilities to compute the expected acceptance length 𝜏. Because neural confidence estimates are often overconfident (Guo et al., 2017; Ovadia et al., 2019), using the raw confidence scores directly would distort the throughput estimation, leading to suboptimal scheduling. To address this, we introduce Sequential Temperature Scaling (STS). Because each 𝑐𝑖 models a conditional probability, the chain rule dictates that the joint probability of a draft prefix being Î accepted factorizes into the cumulative product 𝑖⩽ 𝑘 𝑐𝑖 . Using a held-out validation set, STS calibrates this joint probability consecutively from left to right. Specifically, at each position 𝑘 ∈ {1, . . . , 𝛾 }, we perform a simple 1D grid search to find the optimal temperature scalar that minimizes the Expected Calibration Error (ECE) (Naeini et al., 2015) of the cumulative product, keeping the already-calibrated scores of all preceding positions fixed. Crucially, temperature scaling is an order-preserving transformation: it rectifies the predicted probabilities to match empirical acceptance rates without disrupting the relative draft token rankings learned by the confidence head. 3.2.2. Hardware-Aware Prefix Scheduler 7

在本节中,我们使用离线基准测试验证 DSpark 的草稿质量,并在第 5 节中报告置信度调度器在在线生产流量下的有效性。实验设置在第 4.1 节中描述,主要结果在第 4.2 节中,额外分析包含在第 4.3 节中。4.1. 实验设置 目标模型和草稿模型。我们在四个不同规模和模型系列的目标模型上评估 DSpark:Qwen3-{4B, 8B, 14B}(Yang 等人,2025)和 Gemma4-12B(Google DeepMind,2026)。对于草稿模型,我们将 DSpark 与两个代表性草稿器进行比较:DFlash(Chen 等人,2026),一种最先进的并行草稿器,以及 Eagle3(Li 等人,2026b),一种基于训练时测试(TTT)的自回归草稿器。为了公平比较,我们在相同的训练框架和相同的数据上重新训练所有草稿器。我们将 Eagle3 的 TTT 视野(7)与 DFlash 和 DSpark 使用的块大小(7)对齐,并且对所有草稿器使用相同的目标模型特征层。对于草稿模型层数,我们将 Eagle3 设置为 1,DSpark 和 DFlash 设置为 5(Chen 等人,2026)。除非另有说明,DSpark 表示马尔可夫头变体;我们在第 4.3.2 节中研究 RNN 头变体。训练数据。我们使用 Open-PerfectBlend 2,一个 PerfectBlend(Xu 等人,2024)的开源版本,包含 130 万个样本。它是一个通用指令数据集,包含聊天(17.6%)、数学(39.4%)、代码(38.9%)和指令遵循数据(4.1%)。我们仅使用 Open-PerfectBlend 中的提示;响应由每个目标模型使用推荐的采样参数重新生成。每个草稿器训练 10 个 epoch 以确保完全收敛。对于数据生成和评估,我们采用非思考模式。评估协议。我们在三个领域评估不同算法的性能:1. 数学推理,包括 GSM8K(Cobbe 等人,2021)、MATH500(Lightman 等人,2024)和 AIME25(Zhang 和 Math-AI,2025)。2. 代码生成,包括 MBPP(Austin 等人,2021b)、HumanEval(Chen 等人,2021)和 Live-CodeBench(Jain 等人,2025)。3. 日常聊天,包括 MT-Bench(Zheng 等人,2023)、Alpaca(Taori 等人,2023)和 Arena-Hard(Li 等人,2024a, 2025b)。对于所有基准测试,我们使用标准推测解码(Chen 等人,2023;Leviathan 等人,2023),采样温度设置为 1.0。我们报告每个解码轮次的接受长度(𝜏)3。对于所有草稿器,我们使用基于链的草稿。

Algorithm 1 Hardware-Aware Prefix Scheduler Require: Active requests 𝑟 ∈ {1, . . . , 𝑅 }; confidence sequence 𝑐𝑟,1 , . . . , 𝑐𝑟,𝛾 per request; profiled step curve SPS( 𝐵) Ensure: Selected per-request prefix lengths ℓ1∗ , . . . , ℓ𝑅∗ 1: for 𝑟 = 1 to 𝑅 do Î 2: Compute prefix survival probabilities: 𝑎𝑟, 𝑗 ← 𝑖⩽ 𝑗 𝑐𝑟,𝑖 for 𝑗 = 1, . . . , 𝛾 3: end for 4: Construct candidate space E ← {( 𝑟 , 𝑗) | 𝑎𝑟, 𝑗 > 0} and sort descending by 𝑎𝑟, 𝑗 5: Initialize states: ℓ𝑟 ← 0 for all 𝑟 ; Batch size 𝐵 ← 𝑅; Expected accepts 𝜏∗ ← 𝑅 6: Initialize tracking: Θbest ← 𝑅 · SPS( 𝑅); Selected lengths ℓ𝑟∗ ← 0 for all 𝑟 7: for each ( 𝑟 , 𝑗) ∈ E in sorted order do 8: ℓ𝑟 ← 𝑗; 𝐵 ← 𝐵 + 1; 𝜏∗ ← 𝜏∗ + 𝑎𝑟, 𝑗 9: Current throughput Θ ← 𝜏∗ · SPS( 𝐵) 10: if Θ > Θbest then 11: Θbest ← Θ; Update selected lengths ℓ𝑟∗ ← ℓ𝑟 12: else 13: break 14: end if 15: end for 16: return ( ℓ1∗ , . . . , ℓ𝑅∗ ) achieving Θbest Prior methods (Huang et al., 2024; Li et al., 2024b) typically apply a static threshold to confidence scores to determine verification length. While effective under isolated, single-request assumptions, static thresholds can be suboptimal in high-concurrency production systems, where the utility of verifying a draft token depends heavily on the current system load. To address this, we formulate verification length selection as a global throughput maximization problem (Algorithm 1). Consider a batch of 𝑅 active requests. For request 𝑟 , let 𝑐𝑟,1 , . . . , 𝑐𝑟,𝛾 be the per-position confidence estimates, and let ℓ𝑟 ∈ {0, . . . , 𝛾 } denote the scheduled verification length. Because speculative decoding dynamically accepts draft tokens only as a continuous Î prefix, the survival probability of a token at position 𝑗 is the cumulative product 𝑎𝑟, 𝑗 = 𝑖⩽ 𝑗 𝑐𝑟,𝑖 . In a single verification step, the total batch size (measured in tokens) sent to the target Í model is 𝐵 = 𝑟𝑅=1 (1 + ℓ𝑟 ), and the expected number of successfully accepted tokens is 𝜏 =  Íℓ𝑟 Í𝑅 𝑎 . Let SPS( 𝐵) denote the engine throughput, measured in steps per second, for 𝑟 =1 1 + 𝑗=1 𝑟 , 𝑗 a given forward-pass batch size 𝐵. Crucially, this capacity curve is profiled once during engine initialization and stored as a lightweight cost table. Our scheduler then aims to maximize the expected system-wide token throughput Θ = 𝜏 · SPS( 𝐵) by dynamically selecting verification lengths ℓ1 , . . . , ℓ𝑅 . Although finding the global maximum of Θ appears to be a combinatorial search, the objective structure allows for an efficient greedy solution. Because 𝑎𝑟, 𝑗 is monotonically nonincreasing with respect to 𝑗 (i.e., 𝑎𝑟, 𝑗 ≤ 𝑎𝑟, 𝑗 −1 ), the marginal gain in expected accepted tokens for extending request 𝑟 ’s verification length from 𝑗 − 1 to 𝑗 is exactly 𝑎𝑟, 𝑗 . This monotonicity ensures that sorting candidate tokens globally by 𝑎𝑟, 𝑗 naturally respects intra-block prefix dependencies. Consequently, if the total verification batch size 𝐵 were fixed, the optimal allocation { ℓ𝑟 } would be determined by greedily selecting the draft tokens with the highest survival probabilities from the global pool of all { 𝑎𝑟, 𝑗 }. Building on this insight, the optimization can be evaluated along this greedy admission path. 8

表 1 | 主要推测解码结果。我们报告了不同目标模型和领域的每解码轮次接受长度(𝜏)(越高越好)。粗体表示最佳结果。目标

We first globally sort all valid prefix extensions in descending order of survival probability. To dynamically determine the optimal target batch size 𝐵, we incrementally admit tokens from this sorted pool, updating the expected throughput Θ via an 𝑂 (1) lookup from the pre-profiled cost table. Lossless speculative decoding strictly requires the non-anticipating property: admission decisions must not depend on future candidate tokens (Chen et al., 2023; Leviathan et al., 2023). Because our confidence head relies on the Markov feature of the previously sampled token, computing the next survival probability 𝑎𝑟,𝑘+1 explicitly requires the instantiated candidate 𝑥𝑟,𝑘 . A retrospective global search would thus inadvertently leak 𝑥𝑟,𝑘 into the admission decision for step 𝑘, introducing selection bias (we provide a concrete counterexample demonstrating this theoretical violation in Appendix A). To enforce strict causality, the scheduler (Algorithm 1) employs an early-stopping mechanism. By breaking the greedy search immediately when the throughput drops (Θ ≤ Θbest ), the truncation decision relies solely on the prefix processed up to that exact step. This isolates the admission event from future tokens, ensuring exact target-distribution recovery. Note that this stepwise early-stopping yields the global maximum throughput if and only if the objective Θ is unimodal, which implicitly assumes a smoothly decaying hardware capacity curve. We address the engineering adaptations required for real-world, non-smooth SPS characteristics and asynchronous system pipelines in Section 5.2. 3.3. Training During training, we randomly sample multiple anchor positions from each target sequence to form 𝛾 -token blocks as training data. The target model is frozen throughout training; the draft model shares its embedding layer and language modeling head and keeps them frozen, updating only the backbone drafter, sequential block, and confidence head. The training objective consists of three terms: a cross-entropy loss Lce , a distributionmatching loss Ltv , and a confidence loss Lconf . All three are position-weighted by 𝑤𝑘 = exp(−( 𝑘−1)/𝛾 ) (Chen et al., 2026), which emphasizes earlier block positions that contribute more to the expected acceptance length under prefix-based verification. The cross-entropy loss Lce trains the drafter to predict the correct next token: Lce = −

where 𝑥 𝑘∗ is the ground-truth token and 𝑝𝑘𝑑 is the draft distribution. The distribution-matching loss Ltv penalizes the total variation distance between the draft and target distributions: Ltv =

Since the total variation distance is a direct proxy for the acceptance rate: the per-step acceptance probability equals 1 − 12 ∥ 𝑝𝑑 − 𝑝𝑡 ∥ 1 (Leviathan et al., 2023), minimizing Ltv directly maximizes the expected acceptance rate. The confidence loss Lconf is a binary cross-entropy that trains the confidence head to predict the soft acceptance label 𝑐∗𝑘 from Equation 8: Lconf = −

The overall objective is a weighted combination of the three terms (with default weights 𝛼ce = 0.1, 𝛼tv = 0.9, 𝛼conf = 1.0): L = 𝛼ce Lce + 𝛼tv Ltv + 𝛼conf Lconf (12)

实验 Experiments

In this section, we validate the draft quality of DSpark using offline benchmarks and report the effectiveness of confidence scheduler under online production traffic in Section 5. The experimental setup is described in Section 4.1, main results in Section 4.2, and additional analyses are included in Section 4.3. 4.1. Experimental Setup Target and draft models. We evaluate DSpark on four target models spanning different scales and model families: Qwen3-{4B, 8B, 14B} (Yang et al., 2025), and Gemma4-12B (Google DeepMind, 2026). For draft models, we compare DSpark with two representative drafters: DFlash (Chen et al., 2026), a state-of-the-art parallel drafter, and Eagle3 (Li et al., 2026b), an autoregressive drafter based on Training-Time Test (TTT). For fair comparison, we retrain all drafters in the same training framework and on the same data. We align Eagle3’s TTT horizon (7) with the block size (7) used by DFlash and DSpark, and we use the same target-model feature layers for all drafters. For the number of draft model layers, we set 1 for Eagle3 and 5 for DSpark and DFlash (Chen et al., 2026). Unless otherwise stated, DSpark denotes the Markov-head variant; we study the RNN-head variant in Section 4.3.2. Training data. We use Open-PerfectBlend 2 , an open-sourced version of PerfectBlend (Xu et al., 2024) consisting of 1.3 million samples. It is a general-purpose instruction dataset containing chat (17.6%), math (39.4%), code (38.9%), and instruction-following data (4.1%). We only use the prompts from Open-PerfectBlend; responses are regenerated by each target model with recommended sampling parameters. Each drafter is trained for 10 epochs to ensure full convergence. For data generation and evaluation, we adopt the non-thinking mode. Evaluation protocol. We evaluate the performance of different algorithms on three domains: 1. Mathematical Reasoning, including GSM8K (Cobbe et al., 2021), MATH500 (Lightman et al., 2024) and AIME25 (Zhang and Math-AI, 2025). 2. Code Generation, including MBPP (Austin et al., 2021b), HumanEval (Chen et al., 2021) and Live-CodeBench (Jain et al., 2025). 3. Daily Chat, including MT-Bench (Zheng et al., 2023), Alpaca (Taori et al., 2023) and Arena-Hard (Li et al., 2024a, 2025b). For all benchmarks, we use standard speculative decoding (Chen et al., 2023; Leviathan et al., 2023) with the sampling temperature set to 1.0. We report the accepted length (𝜏) per decoding round3 . For all drafters, we use chain-based drafting.

2 https://huggingface.co/datasets/mlabonne/open-perfectblend 3 For clarity, unless otherwise stated, all reported metrics for accepted length and acceptance rate include the

Table 1 | Main speculative decoding results. We report accepted length (𝜏) per decoding round (higher is better) for different target models and domains. Bold marks the best results. Target

4.2. 实验结果 为了将原始草稿质量与系统级调度策略分离,我们的离线评估禁用了置信度调度器,强制所有草稿模型提出固定数量的令牌。主要结果以每轮平均接受长度(τ)衡量,报告在表 1 中。DSpark 在所有评估的目标模型和基准领域上,始终优于自回归基线(Eagle3)和并行基线(DFlash)。具体来说,在 Qwen3-4B、8B 和 14B 模型上,DSpark 相对于 Eagle3 的宏平均接受长度分别提高了 30.9%、26.7%和 30.0%。类似地,与 DFlash 相比,DSpark 在三个规模上分别实现了 16.3%、18.4%和 18.3%的相对改进。关键的是,这一优势在不同模型家族中具有泛化性,如在 Gemma4-12B 目标上的一致性能提升所示。除了平均改进,表 1 还揭示了强烈的领域效应:在结构化任务(例如,Qwen3-4B 在数学上为 5.57,在代码上为 5.12)上的接受长度自然高于开放式聊天(3.49)。数据可预测性的这种固有方差意味着静态验证长度通常会浪费算力在极有可能被拒绝的尾部令牌上。这直接激发了我们的置信度调度验证,它根据预期接受度动态修剪草稿块。4.3. 实验分析 4.3.1. 为什么并行生成能优于自回归?表 1 呈现了一个反直觉的观察:并行草稿模型(DFlash)和半自回归草稿模型(DSpark)通常比完全自回归草稿模型(Eagle3)产生更长的接受长度。这一发现与标准预期——逐步自回归比并行模型产生更高质量序列——形成对比(Israel et al., 2026; Ren et al., 2020; Zheng et al., 2025)。为了分析这种行为,我们考察了宏级接受长度之外的性能。使用 Qwen3-4B 目标模型和第 4.1 节描述的基准集,我们引入了在实际推测解码过程中跟踪的逐位置条件接受率。具体来说,对于给定的草稿位置 k,评估分母仅计数目标模型成功验证并接受从 1 到 k-1 所有先前草稿令牌的实例。然后,该指标计算在这些有效实例中位置 k 的令牌也被接受的比例。这种方法确保位置 k 的评估不受先前前缀错误的影响,揭示了每个特定步骤的基础预测质量。图 2 详细展示了这些测量结果,显示了不同架构之间的清晰行为差异。位置 1 的容量优势。在第一个草稿位置,两种架构都仅基于目标上下文预测下一个令牌。这里的性能差异严格源于架构容量:像 Eagle3 这样的自回归模型由于其 O(γ)延迟而受限于浅层网络,而 O(1)并行草稿模型可以负担更深的网络。这种结构差距在位置 1 产生了显著的准确率差距,DFlash 的起始点明显高于 Eagle3(例如,数学上 0.88 vs. 0.81,聊天上 0.72 vs. 0.53)。由于推测解码作为严格的前缀匹配生存过程运行,第一个令牌具有最高的杠杆作用——此处的拒绝会立即使整个块无效。因此,这种初始容量优势不成比例地提升了最终接受长度,解释了为什么并行草稿模型尽管在后续位置接受率快速衰减,但最终全局优于自回归模型。后续位置独立性的局限性。检查曲线尾部(位置 2 到 7)暴露了独立并行生成的内在局限性。随着早期令牌锁定特定语义路径,后续令牌自然变得更可预测。像 Eagle3 这样的自回归模型有效利用了这种条件确定性,在块内保持甚至增加条件接受率(例如,在聊天上从 0.53 到 0.74)。相比之下,DFlash 遭受快速接受衰减,在代码上从 0.87 降至 0.78,在聊天上从 0.72 降至 0.63。由于每个并行位置对所有可能的先前令牌进行边缘化,而不是条件于精确采样的前缀,模型经常提出不一致的后缀组合——这种模式称为多模态碰撞(Gu et al., 2018; Stern et al., 2018)。通过半自回归缓解后缀衰减。上述分析突显了一个清晰的架构目标:结合并行骨干网对初始令牌的高容量和自回归模型对后续令牌的依赖建模。这直接

4.2. Experimental Results To isolate the raw draft quality from system-level scheduling policies, our offline evaluation disables the confidence scheduler, forcing all drafters to propose a fixed block of tokens. The main results, measured by the average accepted length (𝜏) per round, are reported in Table 1. DSpark consistently outperforms both the autoregressive baseline (Eagle3) and the parallel baseline (DFlash) across all evaluated target models and benchmark domains. Specifically, across the Qwen3-4B, 8B, and 14B models, DSpark improves the macro-average accepted length over Eagle3 by 30.9%, 26.7%, and 30.0%, respectively. Similarly, compared to DFlash, DSpark yields relative improvements of 16.3%, 18.4%, and 18.3% across the three scales. Crucially, this advantage generalizes across model families, as demonstrated by the consistent performance gains on the Gemma4-12B target. Beyond the average improvements, Table 1 reveals a strong domain effect: the accepted length is naturally higher on structured tasks (e.g., 5.57 on math and 5.12 on code for Qwen3-4B) than on open-ended chat (3.49). This inherent variance in data predictability means a static verification length often wastes compute on trailing tokens that are highly likely to be rejected. This directly motivates our confidence-scheduled verification, which dynamically prunes the draft block based on expected acceptance. 4.3. Experimental Analysis 4.3.1. Why Can Parallel Generation Outperform Autoregression? Table 1 presents a counter-intuitive observation: the parallel drafter (DFlash) and the semiautoregressive drafter (DSpark) often yield longer accepted lengths than the fully autoregressive drafter (Eagle3). This finding contrasts with the standard expectation that step-by-step autoregression produces higher-quality sequences than parallel models (Israel et al., 2026; Ren et al., 2020; Zheng et al., 2025). To analyze this behavior, we examine performance beyond the macro-level accepted length. Using the Qwen3-4B target model and the benchmark sets described in Section 4.1, we introduce position-wise conditional acceptance tracked during actual speculative decoding rollouts. Specifically, for a given draft position 𝑘, the evaluation denominator counts only the instances where 11

图 2 | 逐位置条件接受率。我们报告了每个草稿位置的经验条件接受率,使用 Qwen3-4B 目标模型在每个领域内的基准上取平均。与标准前缀存活率不同,该指标通过移除先前拒绝的惩罚,隔离了位置 k 的基础预测质量。注意,自回归草稿模型(Eagle3)保持稳定或呈上升趋势,而并行草稿模型(DFlash)遭受后缀衰减。

Figure 2 | Position-wise conditional acceptance. We report the empirical conditional acceptance rate for each draft position, averaged across benchmarks within each domain using the Qwen34B target model. Unlike standard prefix survival, this metric isolates the baseline predictive quality at position 𝑘 by removing the penalty of previous rejections. Notice that the autoregressive drafter (Eagle3) remains stable or trends upward, while the parallel drafter (DFlash) suffers suffix decay. the target model successfully verifies and accepts all preceding draft tokens from 1 to 𝑘 − 1. The metric then calculates the proportion of these valid instances where the token at position 𝑘 is also accepted. This approach ensures that the evaluation of position 𝑘 is not penalized by earlier prefix errors, revealing the underlying predictive quality at each specific step. Figure 2 details these measurements, demonstrating clear behavioral differences across the architectures. The Capacity Advantage at Position 1. At the first draft position, both architectures predict the next token based solely on the target context. The performance divergence here stems strictly from architectural capacity: autoregressive models like Eagle3 are constrained to shallow networks due to their 𝑂 ( 𝛾 ) latency, whereas 𝑂 (1) parallel drafters can afford much deeper networks. This structural gap yields a substantial accuracy margin at position 1, with DFlash starting noticeably higher than Eagle3 (e.g., 0.88 vs. 0.81 on Math, and 0.72 vs. 0.53 on Chat). Because speculative decoding operates as a strict prefix-matching survival process, the first token carries the highest leverage—a rejection here immediately invalidates the entire block. Consequently, this initial capacity advantage disproportionately boosts the final accepted length, explaining why parallel drafters ultimately outperform autoregressive ones globally despite rapid acceptance decay at later positions. The Limitation of Independence at Later Positions. Examining the tail of the curves (positions 2 through 7) exposes the inherent limitation of independent parallel generation. As earlier tokens lock in a specific semantic path, subsequent tokens naturally become more predictable. Autoregressive models like Eagle3 effectively leverage this conditional certainty, maintaining or even increasing conditional acceptance deeper into the block (e.g., from 0.53 to 0.74 on Chat). In contrast, DFlash suffers from rapid acceptance decay, dropping from 0.87 to 0.78 on Code and 0.72 to 0.63 on Chat. Because each parallel position marginalizes over all possible prior tokens rather than conditioning on an exact sampled prefix, the model frequently proposes inconsistent suffix combinations—a mode known as multi-modal collision (Gu et al., 2018; Stern et al., 2018). Mitigating Suffix Decay with Semi-Autoregression. The preceding analysis highlights a clear architectural objective: combining the high capacity of a parallel backbone for the initial token with the dependency modeling of an autoregressive model for subsequent tokens. This directly 12

图 3 | 草稿模型深度的影响。在提议长度固定的情况下,DSpark 的性能随着草稿模型层数的增加而提升。值得注意的是,浅层 2 层 DSpark 优于更深的 5 层 DFlash 基线,突显了顺序建模的参数效率。

Figure 3 | Effect of drafter depth. With proposal length fixed, DSpark’s performance improves as drafter layers are added. Notably, a shallow 2-layer DSpark outperforms a deeper 5-layer DFlash baseline, highlighting the parameter efficiency of sequential modeling.

图 4 | 提议长度和延迟开销的影响。DSpark 在各种块大小(左侧三个面板)上始终优于 DFlash。最右侧面板显示,顺序头在服务过程中引入了最小的延迟开销。

Figure 4 | Effect of proposal length and latency overhead. DSpark consistently outperforms DFlash across various block sizes (left three panels). The rightmost panel demonstrates that the sequential head introduces minimal latency overhead during serving. motivates DSpark’s semi-autoregressive design. As shown in Figure 2, DSpark inherits the high initial acceptance of the deep parallel drafter (e.g., starting at 0.93 on Math). Simultaneously, its lightweight sequential head mitigates the rapid acceptance decay typical of parallel generation. By resolving this trade-off, DSpark maintains a high and stable conditional acceptance rate throughout the entire draft block. 4.3.2. A Little Autoregression Goes a Long Way Building on the insights from Section 4.3.1, we explore the architectural design space of DSpark along two dimensions: drafter depth (number of transformer layers) and proposal length (block size 𝛾 ). Unless otherwise stated, all experiments in this section use Qwen3-4B as the target model and follow the evaluation protocol detailed in Section 4.1. Drafter Depth. Increasing the number of transformer layers naturally expands a draft model’s predictive capacity. To isolate this effect, we fix the block size to 7 and vary the number of DSpark layers from 1 to 5, comparing it against a 5-layer DFlash baseline. Figure 3 aggregates the accepted lengths across the math, code, and chat domains. As expected, DSpark’s performance improves monotonically with depth, with the steepest marginal gain occurring from one to two layers. Notably, a 2-layer DSpark outperforms the 5-layer DFlash baseline across all domains. 13

这证明了通过轻量级顺序头注入局部自回归提供了非常有利的准确率-参数权衡,实现了比简单堆叠更深并行层更好的序列连贯性。提议长度。接下来,我们固定草稿模型深度为 5 层,并将草稿长度(提议长度γ加一个锚点令牌)在{4, 8, 12, 16}范围内缩放,以评估在更长草稿块上的性能。对于 DSpark,我们评估了默认的马尔可夫头和 RNN 头。图 4 的前三个面板显示,DSpark 在每个提议长度上始终优于 DFlash。更重要的是,随着γ增加,性能差距稳步扩大。由于纯并行生成(DFlash)遭受快速接受衰减(图 2),其边际效用对于长块会减小。DSpark 缓解了这种衰减,导致其相对于 DFlash 的相对增益增长。例如,在γ=7 时,DSpark 在数学上提高了 16%的接受长度,在代码上提高了 15%,在聊天上提高了 18%;在γ=15 时,这些增益分别扩大到 30%、26%和 22%。此外,RNN 头相对于马尔可夫头仅提供边际额外增益,主要是在更长的提议长度上。鉴于其更高的实现复杂性和不太有利的部署特性,我们使用马尔可夫头作为默认。延迟开销。我们量化了 DSpark 中顺序生成循环的开销。图 4 最右侧面板报告了每轮引擎延迟——包括一次目标验证过程、并行草稿块前向和串行采样循环——在批大小为 128 时测量。为了防止序列长度偏差,报告的延迟是不同上下文长度({512, 1024, 2048, 4096}令牌)的算术平均值。由于在此批大小下目标模型主导验证计算时间,顺序块的延迟开销可以忽略不计。因此,将草稿长度从 4 扩展到 16,相对于 DFlash 基线,仅增加了 0.2%到 1.3%的全轮延迟,尽管接受长度提高了高达 30%。4.3.3. 更智能而非更长的验证:置信度头的作用虽然 DSpark 在长草稿块上维持高接受率,但验证整个提议仍然低效(Hu et al., 2026; Huang et al., 2024)。由于第 4.2 节提到的固有领域方差,开放式聊天中的尾部令牌仍然面临高拒绝风险,使得盲目验证浪费目标算力。为了评估置信度头能否有效修剪这些无前途的后缀,我们使用 Qwen3-4B 进行了离线阈值扫描。我们在此单独验证估计器,将硬件感知前缀调度器(第 3.2.2 节)留待第 5 节进行在线生产评估。诊断:静态阈值扫描。图 5 绘制了不同置信度阈值下的每步平均令牌数(柱状图)和整体接受率(线图)。随着阈值增加,接受率稳步上升,因为估计器过滤掉了最终会被拒绝的令牌(阴影柱)。这表明置信度头能够识别低价值后缀令牌,并且这种修剪在聊天工作负载上最为显著,其中高熵令牌分布限制了固定长度验证的效率。在聊天子图中,提高阈值显著减少了被拒绝的令牌,将接受率从 45.7%提高到 95.7%。相比之下,结构化任务(数学和代码)经历较温和的修剪并保留更多草稿令牌,接受率分别从 76.9%提高到 92.5%,从 67.6%提高到 92.0%。

This demonstrates that injecting local auto-regression via a lightweight sequential head offers a highly favorable accuracy-parameter trade-off, achieving better sequence coherence than simply stacking deeper parallel layers. Proposal Length. Next, we fix the drafter depth to 5 layers and scale the draft length (proposal length 𝛾 plus one anchor token) across {4, 8, 12, 16} to evaluate performance on longer draft blocks. For DSpark, we evaluate both the default Markov head and the RNN head. The first three panels of Figure 4 show that DSpark consistently outperforms DFlash at every proposal length. More importantly, the performance gap steadily widens as 𝛾 increases. Because pure parallel generation (DFlash) suffers from rapid acceptance decay (Figure 2), its marginal utility diminishes for long blocks. DSpark mitigates this decay, causing its relative gain over DFlash to grow. For instance, at 𝛾 = 7, DSpark improves the accepted length by 16% on math, 15% on code, and 18% on chat; at 𝛾 = 15, these gains expand to 30%, 26%, and 22%, respectively. Also, RNN head provides only marginal additional gains over the Markov head, mainly at longer proposal lengths. Given its higher implementation complexity and less favorable deployment properties, we use the Markov head as the default. Latency Overhead. We quantify the overhead of the sequential generation loop in DSpark. The rightmost panel of Figure 4 reports the per-round engine latency—comprising one target verification pass, the parallel draft block forward, and the serial sampling loop—measured at a batch size of 128. To prevent sequence-length bias, the reported latency represents the arithmetic mean across varying context lengths ({512, 1024, 2048, 4096} tokens). Since the target model dominates the verification compute time at this batch size, the sequential block’s latency overhead is negligible. Consequently, scaling the draft length from 4 to 16 adds a marginal 0.2% to 1.3% to the full-round latency over the DFlash baseline, despite delivering up to a 30% improvement in accepted length. 4.3.3. Verify Smarter, Not Longer: The Role of Confidence Head While DSpark sustains high acceptance over long draft blocks, verifying the entire proposal remains inefficient (Hu et al., 2026; Huang et al., 2024). Due to the inherent domain variance noted in Section 4.2, trailing tokens in open-ended chat still face high rejection risks, making blind verification a waste of target compute. To evaluate whether the confidence head can effectively prune these unpromising suffixes, we conduct an offline threshold sweep using Qwen3-4B. We validate the estimator in isolation here, reserving the hardware-aware prefix scheduler (Section 3.2.2) for live production evaluation in Section 5. Diagnostic: Static Threshold Sweep. Figure 5 plots the average tokens per step (bars) and the overall acceptance rate (line) across confidence thresholds. As the threshold increases, the acceptance rate steadily rises because the estimator filters out tokens that would ultimately be rejected (hashed bars). This suggests that the confidence head can identify lower-value suffix tokens and this pruning is most pronounced on chat workloads, where higher-entropy token distributions limit the efficiency of fixed-length verification. In the Chat subplot, raising the threshold significantly reduces rejected tokens, increasing the acceptance rate from 45.7% to 95.7%. In contrast, structured tasks (Math and Code) experience milder pruning and retain more draft tokens, with acceptance rates rising from 76.9% to 92.5% and 67.6% to 92.0%, respectively.

图 5 | 置信度阈值扫描。阈值为 0 对应标准固定长度验证。随着阈值增加,整体接受率稳步上升,因为置信度头有效修剪了最终会被拒绝的令牌(阴影柱)。完美校准

Figure 5 | Confidence threshold sweep. A threshold of 0 corresponds to standard fixed-length verification. As the threshold increases, the overall acceptance rate steadily rises because the confidence head effectively prunes tokens that would ultimately be rejected (hashed bars). Perfect calibration

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

图 6 | Alpaca 数据集上的可靠性图。虽然原始置信度估计器实现了强判别能力,但其预测固有地过于自信。应用事后校准有助于使前缀存活概率与经验接受率对齐。阴影背景直方图表示不同置信度区间内样本计数的频率分布。从静态阈值到校准调度。虽然对诊断有用,但静态阈值在动态服务环境中是次优的,因为它忽略了系统负载:在低并发下验证低置信度令牌的机会成本最小,但在高并发下会浪费关键批容量。这种负载依赖性激发了硬件感知前缀调度器。如第 3.2 节所述,最大化系统级吞吐量要求置信度模型既具有强预测判别能力,又具有精确校准以准确估计累积存活概率。可靠性图(图 6)表明,虽然原始模型实现了强判别能力(ROC-AUC(Hanley and McNeil, 1982)范围从 0.81 到 0.90),但它过于自信(ECE 3%–8%)。应用事后 STS(第 3.2.1 节)缓解了这种过度自信,将平均 ECE 降低到约 1%,并产生可靠的存活估计。

Figure 6 | The Reliability Diagram on Alpaca Dataset. While the raw confidence estimator achieves strong discrimination, its predictions are inherently overconfident. Applying post-hoc calibration helps to align the prefix survival probabilities with empirical acceptance rates. The shaded background histogram represents the frequency distribution of sample counts across different confidence bins. From Static Thresholds to Calibrated Scheduling. While useful for diagnostics, a static threshold is sub-optimal in dynamic serving environments because it ignores system load: verifying low-confidence tokens incurs minimal opportunity cost under low concurrency, but wastes critical batch capacity under high concurrency. This load dependency motivates the hardware-aware prefix scheduler. As formulated in Section 3.2, maximizing system-level throughput requires the confidence model to exhibit both strong predictive discrimination and precise calibration to accurately estimate cumulative survival probabilities. The reliability diagram (Figure 6) demonstrates that while the raw model achieves strong discrimination (ROC-AUC (Hanley and McNeil, 1982) ranging from 0.81 to 0.90), it is overly confident (ECE 3%–8%). Applying post-hoc STS (Section 3.2.1) mitigates this overconfidence, reducing the average ECE to ∼1% and yielding reliable survival estimates.

实际部署 Real-World Deployment

虽然第 4 节确立了 DSpark 在离线基准上的算法增益,但将其与 DeepSeek-V4(DeepSeek-AI, 2026)等大规模模型一起部署引入了额外的

While Section 4 establishes the algorithmic gains of DSpark on offline benchmarks, deploying it alongside large-scale models like DeepSeek-V4 (DeepSeek-AI, 2026) introduces additional 15

系统级挑战,涉及训练和推理。在本节中,我们介绍 DSpark 的端到端生产流水线。我们详细介绍了可扩展的训练机制、部署硬件感知前缀调度器(第 3.2.2 节)所需的系统级优化,以及框架在实时用户流量下的端到端性能。5.1. 可扩展且灵活的训练 DSpark 草稿模型与 DeepSeek-V4-Flash 和 DeepSeek-V4-Pro(DeepSeek-AI, 2026)的预览版本共同部署。并行骨干网包含三个 MoE 层(Dai et al., 2024),采用 mHC(Xie et al., 2026)和滑动窗口注意力(窗口大小为 128)。我们将最大块大小配置为γ=5,并使用马尔可夫头进行顺序建模。此外,置信度头与草稿模型一起端到端训练,随后通过 STS 进行校准,以提供可靠的调度信号。训练草稿模型需要目标模型的输出分布作为监督。在完整文档上下文中评估两个模型会带来显著的内存占用和工作者间通信开销。为了解决这些瓶颈,我们在内部训练框架(HAI-LLM)中实现了两个系统级优化:• 隐藏状态通信。在并行工作者之间传输目标模型的全词汇 logits(V≈10^5)会造成显著的带宽瓶颈。相反,我们临时缓存目标模型前向传播的激活,并仅传递紧邻语言建模(LM)头的隐藏状态。然后,仅在草稿模型工作者上对采样的目标位置本地执行 LM 头投影。这将每令牌通信复杂度降低到 O(d),其中 d 是隐藏维度。• 锚点边界序列打包。为了将草稿模型的计算成本与目标模型的上下文长度解耦,我们从训练序列中采样固定数量的草稿锚点,并将这些独立的预测块打包成密集的训练批次。我们通过令牌级注意力索引而不是标准 2D 掩码来管理这种打包。这保持了跨多个独立序列和锚点的精确因果掩码,避免了与标准填充相关的计算和内存开销。5.2. 硬件感知前缀调度器在实践中 在第 3.2.2 节中,算法 1 提供了一种理论上合理且无损的调度机制。然而,直接将该算法部署到生产环境中会暴露出两个与现实基础设施的基本冲突。首先,该算法假设平滑的单峰容量曲线,而实际硬件容量 SPS(B)固有地是离散的,表现出锯齿状的阶梯式退化(Yan et al., 2020)。其次,该算法要求每步调度动态草稿令牌,这与连续的 CUDA 图重放(Fireworks AI, 2023)和零开销调度(ZOS)(Zheng et al., 2024; Zhu et al., 2025)相冲突。为了在系统兼容性、吞吐量和算法正确性之间进行权衡,我们调整调度器以异步方式运行。由于 ZOS 要求下一步的批大小在当前步完成之前已知,同步调度将不可避免地导致 GPU 流水线停顿。相反,我们使用两步前的置信度头输出来近似即将到来的验证容量。在机制上,当前步中的候选令牌仍然严格按其实际、最新的累积置信度分数排序;两步前的历史预测仅用于确定动态截断长度(即批容量限制 K)。这实际上将准入过程转化为动态 top-K 选择。虽然近似容量 K 引入了轻微的时间偏移,但选择机制本质上是保序的:最自信的草稿令牌总是优先被验证。这种适应完全隐藏了调度延迟,并确保与 ZOS 无缝集成。基于这种异步流水线,我们解决了硬件利用率的瓶颈。为了防止调度器被锯齿状的 SPS 悬崖困在局部最小值中,我们移除了提前停止的 break,允许无约束的全局搜索。通常,这种回溯搜索会泄露未来令牌信息并违反无损保证(附录 A)。然而,我们的 ZOS 驱动适应自然防止了这一点。因为无约束搜索仅评估两步前的历史预测,准入决策与当前令牌 x_{r,k}的实现隔离。截断长度本质上仅依赖于两步前可用的信息。因此,异步设计形成了因果屏障,在硬件悬崖上最大化物理吞吐量,同时保持精确的目标分布。5.3. 高吞吐量和低延迟推理 在解码过程中,生产服务系统必须同时优化两个竞争目标:每请求延迟和聚合吞吐量(Kwon et al., 2023; Zhao et al., 2025a; Zhong et al., 2024)。前者决定单个用户的服务质量——这一因素在基于智能体的工作负载中日益关键(Tiwari et al., 2026)——而后者决定同时服务的用户总数。由于推测解码不可避免地会浪费验证算力,它固有地处理这种权衡,用额外的系统算力换取更快的每请求生成。然而,在我们的部署环境中,每步处理的请求数量经常受到资源限制(例如,每请求固定的 KV 缓存容量)和可用用户流量池(例如,强化学习长尾负载)的约束。因此,有效批大小始终远低于 GPU 的算力饱和阈值。在这种机制下,传统的权衡简化了:给定固定的并发限制,最大化每 GPU 总令牌吞吐量和最大化每用户生成速度(tok/s/user)成为高度相关的目标,而不是竞争目标。为了实现这种最大吞吐量,异步调度器(第 5.2 节)主动将空闲算力路由到最有希望的草稿令牌。然而,执行这种动态路由在物理执行层引入了严峻挑战:推理框架必须高效支持单批内的可变长度查询。标准解码内核针对固定查询长度进行了高度优化;天真地处理可变长度验证前缀会导致严重的 GPU 利用率不足,由于填充和不均匀的工作负载分布。我们通过将物理执行与逻辑序列跟踪解耦来解决这个问题。在我们的计算内核中,不同请求的所有令牌被展平并作为独立元素相同处理。复杂的序列内依赖关系则通过集成到稀疏注意力实现中的标记张量严格传达。具体在 DeepSeek-V4 架构上,只有索引注意力和压缩内核需要修改以支持这种可变长度路由,使动态调度器能够无缝运行,而不会引入低级执行开销。

system-level challenges across both training and inference. In this section, we present the end-toend production pipeline of DSpark. We detail our scalable training mechanisms, the system-level optimizations necessary to deploy the hardware-aware prefix scheduler (Section 3.2.2), and the framework’s end-to-end performance under live user traffic. 5.1. Scalable and Flexible Training The DSpark draft models are co-deployed with the preview versions of DeepSeek-V4-Flash and DeepSeek-V4-Pro (DeepSeek-AI, 2026). The parallel backbone comprises three MoE layers (Dai et al., 2024) with mHC (Xie et al., 2026) and a sliding window attention of 128. We configure the maximum block size to 𝛾 = 5 and utilize the Markov head for sequential modeling. Furthermore, the confidence head is trained end-to-end alongside the draft model and subsequently calibrated via STS to provide reliable scheduling signals. Training the draft model requires the target model’s output distributions for supervision. Evaluating both models over the full document context incurs substantial memory footprints and inter-worker communication overhead. To address these bottlenecks, we implement two system-level optimizations within our internal training framework (HAI-LLM)4 : • Hidden state communication. Transferring the target model’s full-vocabulary logits (𝑉 ≈ 105 ) across parallel workers creates a significant bandwidth bottleneck. Instead, we temporarily cache the target model’s forward-pass activations and communicate only the hidden states immediately preceding the language modeling (LM) head. The LM head projection is then executed locally on the draft model’s workers only for the sampled target positions. This reduces the per-token communication complexity to 𝑂 ( 𝑑 ), where 𝑑 is the hidden dimension. • Anchor-bounded sequence packing. To decouple the draft model’s computational cost from the target model’s context length, we sample a fixed number of draft anchors from the training sequence and pack these isolated prediction blocks into dense training batches. We manage this packing via token-level attention indices rather than standard 2D masks. This maintains exact causal masking across multiple independent sequences and anchors, avoiding the computational and memory overhead associated with standard padding. 5.2. Hardware-Aware Prefix Scheduler in Practice In Section 3.2.2, Algorithm 1 provides a theoretically sound and lossless scheduling mechanism. However, directly deploying this algorithm into a production environment exposes two fundamental conflicts with real-world infrastructure. First, the algorithm assumes a smooth, unimodal capacity curve, whereas the true hardware capacity SPS( 𝐵) is inherently discrete, exhibiting a jagged, step-wise degradation (Yan et al., 2020). Second, the algorithm requires scheduling of dynamic draft tokens per step, which clashes with continuous CUDA graph replay (Fireworks AI, 2023) and Zero-Overhead Scheduling (ZOS) (Zheng et al., 2024; Zhu et al., 2025). To navigate the trade-offs among system compatibility, throughput, and algorithmic correctness, we adapt the scheduler to operate asynchronously. Because ZOS requires the batch size for the next step to be known before the current step completes, synchronous scheduling would inevitably stall the GPU pipeline. Instead, we approximate the upcoming verification capacity using the confidence head outputs from two steps prior. Mechanically, the candidate tokens in the current step are still strictly sorted by their actual, up-to-date cumulative confidence scores; the historical prediction from two steps prior is used solely to determine the dynamic truncation 4 https://www.high-flyer.cn/en/blog/hai-llm/

TPS (token/s/user) 图 7 | 吞吐量 vs. TPS。在实时流量下,聚合输出令牌吞吐量与每请求生成速度(tok/s/user)的关系。在我们的生产部署中,DSpark 相对于 MTP-1 基线,在测量的流量和引擎配置下改善了观察到的吞吐量-交互性前沿。5.4. 实时用户流量下的性能 我们评估了 DSpark-5(配置最大草稿长度γ=5)与 MTP-1(DeepSeek-AI, 2024)基线在 DeepSeek-V4-Flash(预览版)和 DeepSeek-V4-Pro(预览版)的生产服务引擎中的性能。MTP-1 代表了先前的生产设置,在 DeepSeek-V4-preview 发布两周后被 DSpark 取代。这种单令牌设置历史上在生产中维护,因为部署静态多令牌草稿模型(例如,MTP-3/5)在高并发下由于过多的验证开销会严格降低聚合吞吐量。因此,将 DSpark 与此既定基线直接比较,直接证明了其在动态服务环境中安全解锁更大草稿块性能潜力的能力。在所有图中,散点代表直接从实时用户流量采样的原始遥测数据,捕获了复杂的真实世界请求分布,而实线代表拟合的性能前沿。服务帕累托前沿。图 7 展示了系统聚合吞吐量与每用户生成速度(交互性)之间的权衡。为了量化 DSpark 在实际部署约束下的行为,我们在几个交互性 SLA 锚点评估系统。这里,SLA(服务等级协议)指定了系统必须保证的最小每用户生成速度(以每秒令牌数计)。对于 V4-Flash 引擎,我们在 80 和 120 tok/s/user 的 SLA 锚点评估系统。在适中的 80 tok/s/user SLA 下,DSpark 相对于 MTP-1 基线将聚合吞吐量提高了 51%。更严格的 120 tok/s/user SLA 代表了一个定性不同的机制:在此约束下,单令牌 MTP-1 基线接近其操作边界,只能维持非常小的并发批。因此,此点的相对吞吐量比率数值很大,DSpark 实现了名义上 661%更高的聚合吞吐量。因此,我们主要将此高 SLA 点解释为 DSpark 扩展了可行的交互性前沿的证据,而不是相对于充分利用的基线的代表性加速倍数。在匹配的实际吞吐量水平下(提供更稳定的比较),DSpark 将每用户生成速度提高了 60%到 85%。

length (i.e., the batch capacity limit 𝐾 ). This effectively casts the admission process as a dynamic top- 𝐾 selection. While approximating the capacity 𝐾 introduces a slight temporal offset, the selection mechanism is fundamentally rank-preserving: the most confident draft tokens are always prioritized for verification. This adaptation fully hides scheduling latency and ensures seamless ZOS integration. Building on this asynchronous pipeline, we resolve the hardware utilization bottleneck. To prevent the scheduler from being trapped in local minima by jagged SPS cliffs, we remove the early-stopping break, enabling an unconstrained global search. Ordinarily, this retrospective search would leak future token information and violate the lossless guarantee (Appendix A). However, our ZOS-driven adaptation naturally prevents this. Because the unconstrained search evaluates only historical predictions from two steps prior, the admission decision is isolated from the realization of the current token 𝑥𝑟,𝑘 . The truncation length inherently depends only on information available from two steps prior. Thus, asynchronous design forms a causal barrier, maximizing physical throughput across hardware cliffs while preserving the exact target distribution. 5.3. High-Throughput and Low-Latency Inference During decoding, production serving systems must simultaneously optimize two competing objectives: per-request latency and aggregate throughput (Kwon et al., 2023; Zhao et al., 2025a; Zhong et al., 2024). The former governs the quality of service for individual users—a factor increasingly critical in agent-based workloads (Tiwari et al., 2026)—while the latter determines the total number of concurrently served users. Because speculative decoding inevitably incurs wasted verification compute, it inherently navigates this trade-off, trading extra system compute for faster per-request generation. In our deployment setting, however, the number of requests processed per step is frequently constrained by resource limits (e.g., fixed KV-cache capacity per request) and the pool of available user traffic (e.g., RL long-tail loads). Consequently, the effective batch size persistently remains well below the GPU’s compute-saturating threshold. Under this regime, the traditional trade-off simplifies: given a fixed concurrency limit, maximizing per-GPU total token throughput and maximizing the generation speed per user (tok/s/user) become highly correlated objectives rather than competing ones. To achieve this maximum throughput, the asynchronous scheduler (Section 5.2) actively routes idle compute toward the most promising draft tokens. However, executing this dynamic routing introduces a severe challenge at the physical execution layer: the inference framework must efficiently support variable-length queries within a single batch. Standard decode kernels are heavily optimized for fixed query lengths; naively processing variable-length verified prefixes leads to severe GPU under-utilization due to padding and uneven workload distribution. We resolve this by decoupling physical execution from logical sequence tracking. In our compute kernels, all tokens across different requests are flattened and processed identically as independent elements. The complex intra-sequence dependencies are then strictly conveyed via a marker tensor integrated into our sparse attention implementation. Specifically on the DeepSeek-V4 architecture, only the index-attention and compress kernels require modification to support this variable-length routing, allowing the dynamic scheduler to operate seamlessly without introducing low-level execution overhead.

V4-Pro 部署显示出相同的模式。在适中的 35 tok/s/user SLA 下,DSpark 将聚合吞吐量提高了 52%。在更严格的 50 tok/s/user SLA 下,MTP-1 再次进入低并发机制,DSpark 获得名义上 406%的相对吞吐量优势。与 V4-Flash 一样,我们将此点视为 DSpark 在基线无法有效支持的交互性目标下维持有用吞吐量的指示。在匹配的系统容量下,DSpark 提供 57%到 78%更快的每用户生成。总体而言,这些结果表明 DSpark 将观察到的吞吐量-交互性前沿向外移动:它在适中 SLA 机制下提高了吞吐量,更重要的是,在严格的交互性约束下保持了非退化的服务容量。

TPS (token/s/user) Figure 7 | Throughput vs. TPS. Aggregate output token throughput against per-request generation speed (tok/s/user) under live traffic. In our production deployment, DSpark improves the observed throughput–interactivity frontier relative to the MTP-1 baseline under the measured traffic and engine configurations. 5.4. Performance under Live User Traffic We evaluate DSpark-5 (configured with a maximum draft length of 𝛾 = 5) against the MTP1 (DeepSeek-AI, 2024) baseline within the production serving engines of DeepSeek-V4-Flash (preview) and DeepSeek-V4-Pro (preview). MTP-1 represents the former production setup, having been superseded by DSpark two weeks following the DeepSeek-V4-preview release. This singletoken setup was historically maintained in production because deploying a static multi-token drafter (e.g., MTP-3/5) strictly degrades aggregate throughput under high concurrency due to excessive verification overhead. Therefore, comparing DSpark against this established baseline directly demonstrates its ability to safely unlock the performance potential of larger draft blocks in dynamic serving environments. In all figures, the scatter points represent raw telemetry data sampled directly from live user traffic, capturing complex, real-world request distributions, while the solid lines represent the fitted performance frontiers. The Serving Pareto Frontier. Figure 7 illustrates the trade-off between aggregate system throughput and per-user generation speed (interactivity). To quantify DSpark’s behavior under practical deployment constraints, we evaluate the system at several interactivity SLA anchors. Here, an SLA (Service Level Agreement) specifies the minimum per-user generation speed (in tokens per second) that the system must guarantee. For the V4-Flash engine, we evaluate the system at SLA anchors of 80 and 120 tok/s/user. At the moderate 80 tok/s/user SLA, DSpark improves aggregate throughput by 51% over the MTP-1 baseline. The stricter 120 tok/s/user SLA represents a qualitatively different regime: under this constraint, the single-token MTP-1 baseline approaches its operational boundary and can sustain only a very small concurrent batch. Consequently, the relative throughput ratio at this point is numerically large, with DSpark achieving a nominal 661% higher aggregate throughput. We therefore interpret this high-SLA point primarily as evidence that DSpark extends the feasible interactivity frontier, rather than as a representative multiplicative speedup over a well-utilized baseline. At matched practical throughput levels, which provide a more stable comparison, DSpark accelerates per-user generation speeds by 60% to 85%.

图 8 | 负载自适应吞吐量和验证预算。顶行(a, b):不同系统并发水平下的聚合输出吞吐量。底行(c, d):每请求分配的平均目标验证预算。随着并发负载增加,动态调度器自动限制每请求验证长度以防止资源争用。负载下的吞吐量动态。图 8 通过绘制聚合吞吐量(顶行)和动态验证预算(底行)与系统并发度的关系,分析了驱动这些增益的底层机制。• 在生产部署典型的适中并发机制下(V4-Flash 少于 200 个并发请求,V4-Pro 少于 150 个),硬件感知调度器通过分配更长的验证预算来利用可用的目标算力容量,从 MTP-1 的静态 2 个令牌扩展到每请求大约 4-6 个令牌。这种扩展的验证在每个前向传播中产生更多接受的令牌,直接贡献了在帕累托前沿观察到的吞吐量增益。• 随着系统并发度扩展且目标容量饱和,调度器动态限制此预算。平均验证长度随负载平滑减小,确保低置信度草稿令牌在消耗关键批容量之前被修剪。这种负载感知行为稳定了生产部署:DSpark 在轻流量下最大化空闲算力的效用,同时在重流量下有效保留关键批容量。

The V4-Pro deployment shows the same pattern. At the moderate 35 tok/s/user SLA, DSpark improves aggregate throughput by 52%. At the stricter 50 tok/s/user SLA, MTP-1 again enters a low-concurrency regime, yielding a nominal 406% relative throughput advantage for DSpark. As with V4-Flash, we treat this point as an indication that DSpark sustains useful throughput under an interactivity target that the baseline cannot efficiently support. At matched system capacities, DSpark delivers 57% to 78% faster per-user generation. Overall, these results show that DSpark shifts the observed throughput–interactivity frontier outward: it improves throughput in moderate-SLA regimes and, more importantly, preserves non-degenerate serving capacity under strict interactivity constraints.

局限性。尽管前缀调度器最小化了浪费的目标模型验证,但 DSpark 仍然通过并行骨干网生成初始γ令牌块时产生固定的草稿侧成本。对于固有接受率低的复杂查询,这种前期草稿算力是不可恢复的。未来的优化可以在草稿模型中引入难度感知的提前退出,使此类请求能够绕过全块生成。

Figure 8 | Load-adaptive throughput and verification budgets. Top row (a, b): Aggregate output throughput across varying levels of system concurrency. Bottom row (c, d): The average target verification budget allocated per request. As concurrent load increases, the dynamic scheduler automatically restricts the per-request verification length to prevent resource contention. Throughput Dynamics under Load. Figure 8 analyzes the underlying mechanism driving these gains by plotting aggregate throughput (top row) and the dynamic verification budget (bottom row) against system concurrency. • Under the moderate concurrency regimes typical of our production deployment (fewer than 200 concurrent requests for V4-Flash and 150 for V4-Pro), the hardware-aware scheduler leverages available target compute capacity by allocating longer verification budgets, expanding from MTP-1’s static 2 tokens to roughly 4–6 tokens per request. This extended verification yields more accepted tokens per forward pass, directly contributing to the throughput gains observed on the Pareto frontier. • As system concurrency scales and target capacity saturates, the scheduler dynamically restricts this budget. The average verification length decreases smoothly with load, ensuring that low-confidence draft tokens are pruned before they consume critical batch capacity. This load-aware behavior stabilizes production deployment: DSpark maximizes

推测解码算法。推测解码通过将令牌提议与验证解耦来加速自回归生成。基于早期的块级方法(Ge et al., 2022; Stern et al., 2018; Sun et al., 2021; Xia et al., 2023),现代方法采用拒绝采样来精确保持目标模型的分布(Chen et al., 2023; Leviathan et al., 2023)。由于推理加速直接取决于草稿模型的效率和准确性,大量研究集中在优化其架构上。除了使用独立的语言模型(Chen et al., 2023; Leviathan et al., 2023),后续工作将多头令牌或特征外推器直接集成到目标模型中(Ankner et al., 2024; Cai et al., 2024, 2025; DeepSeek-AI, 2024; Gloeckle et al., 2024; Li et al., 2024b,c, 2026b; Zhang et al., 2025)。其他策略包括通过提前退出进行自我推测(Elhoushi et al., 2024; Liu et al., 2024a; Xia et al., 2025; Zhang et al., 2024)、动态词汇压缩(Williams et al., 2026; Zhao et al., 2025b)、提示查找(Saxena, 2023; Somasundaram et al., 2025)、后缀自动机(Hu et al., 2025)和检索(He et al., 2023; Shen et al., 2026)。为了消除草稿本身的顺序瓶颈,最近的方法提出了并行或块级生成。P-EAGLE 并行化了 EAGLE 风格的草稿(Hui et al., 2026),而 PARD、DART 和 DFlash 使用扩散启发的预测在单个前向传播中生成整个块(An et al., 2026; Chen et al., 2026; Liu et al., 2026a),DDTree 随后将其扩展为可验证的草稿树(Ringel and Romano, 2026)。同时的努力也改进了 DFlash:Domino(Huang et al., 2026a)引入了因果编码器,概念上类似于我们的 RNN 头,而 DFlare(Zhang et al., 2026a)通过逐层融合解决了条件瓶颈。推测解码的系统感知调度。除了草稿模型架构,另一条工作线专注于确定每轮生成或验证的最佳推测令牌数量。为此,各种方法使用置信度启发式(Du et al., 2024; Li et al., 2024b; Liu et al., 2026c; Mamou et al., 2024; Wen and Feng, 2026)、学习到的接受预测器(Huang et al., 2024; Zacks917, 2026)或赌博机策略(Liu et al., 2026b)动态调整草稿长度。此外,认识到推测解码本质上是一个系统级调度问题,最近的工作通过根据实时系统负载和请求优先级调整推测预算来优化整体有效吞吐量和延迟(AngelSlim Team, 2026; Hu et al., 2026; Huang et al., 2026b; Li et al., 2026a; Liu et al., 2024c; Miao et al., 2024; Sadhukhan et al., 2025; Wu et al., 2025)。并行生成。并行生成令牌的模型提供几乎与输出长度无关的解码延迟,使其成为自回归解码的有吸引力的替代方案。非自回归 Transformer(NATs, Gu et al., 2018)通过单次传播独立预测所有位置开创了这一方向。然而,这迫使模型对所有

the utility of idle compute under light traffic, while effectively preserving critical batch capacity under heavy traffic. Limitations. Although the prefix scheduler minimizes wasted target-model verification, DSpark still incurs a fixed draft-side cost to generate the initial 𝛾 -token block via the parallel backbone. For complex queries with inherently low acceptance rates, this upfront drafting compute is unrecoverable. Future optimizations could introduce difficulty-aware early exiting within the draft model, enabling such requests to bypass full-block generation.

相关工作 Related Work

合理模式进行平均,通常产生混合来自不同有效序列片段的输出。已经出现了两条广泛的研究线来解决这一局限性。一条方向保留单次传播架构,但改变模型看到的内容或训练方式:引入潜在变量作为条件输入以将所有位置引导向一致输出(Gu et al., 2018; Kaiser et al., 2018; Ma et al., 2019),或放宽训练目标使模型专注于产生单个连贯输出,而不是对所有有效替代方案的完整分布进行建模(Du et al., 2021; Qian et al., 2021; Shao et al., 2021, 2023)。另一条方向通过迭代重新预测(Austin et al., 2021a; Ghazvininejad et al., 2019; Li et al., 2022)、块级自回归(Arriola et al., 2025; Wang et al., 2018)或结构化输出层(如 CRF(Sun et al., 2019)、CTC(Libovický and Helcl, 2018; Saharia et al., 2020)、HMM(Huang et al., 2022b)和 PCFG(Gui et al., 2023))重新引入有限的顺序依赖性。推测解码进一步要求草稿模型必须为拒绝采样规则提供精确的逐令牌概率。由于迭代细化、潜在边缘化或全局归一化,上述大多数技术无法轻易提供此类概率。例如,在与我们设计密切相关的 CRF-NAT(Sun et al., 2019)中,也在并行隐藏状态上放置了顺序模块,但其全局归一化的配分函数阻止了精确的逐令牌概率计算。类似地,当将 CTC 输出层适应于并行推测解码时,CTC-drafter(Wen et al., 2024)由于对齐路径的潜在边缘化而局限于贪婪验证。DSpark 通过保持顺序校正局部化来规避这些限制,因此逐令牌概率仍然是精确的 softmax 评估。

Speculative Decoding Algorithms. Speculative decoding accelerates autoregressive generation by decoupling token proposal from verification. Building on early blockwise methods (Ge et al., 2022; Stern et al., 2018; Sun et al., 2021; Xia et al., 2023), modern approaches employ rejection sampling to exactly preserve the target model’s distribution (Chen et al., 2023; Leviathan et al., 2023). Because inference speedup directly depends on the drafter’s efficiency and accuracy, extensive research has focused on optimizing its architecture. Beyond using standalone small language models (Chen et al., 2023; Leviathan et al., 2023), subsequent work integrates multi-token heads or feature extrapolators directly into the target model (Ankner et al., 2024; Cai et al., 2024, 2025; DeepSeek-AI, 2024; Gloeckle et al., 2024; Li et al., 2024b,c, 2026b; Zhang et al., 2025). Other strategies include self-speculation via early exits (Elhoushi et al., 2024; Liu et al., 2024a; Xia et al., 2025; Zhang et al., 2024), dynamic vocabulary compression (Williams et al., 2026; Zhao et al., 2025b), prompt lookup (Saxena, 2023; Somasundaram et al., 2025), suffix automata (Hu et al., 2025), and retrieval (He et al., 2023; Shen et al., 2026). To remove the sequential bottleneck of drafting itself, recent methods propose parallel or blockwise generation. P-EAGLE parallelizes EAGLE-style drafting (Hui et al., 2026), while PARD, DART, and DFlash use diffusion-inspired prediction to generate entire blocks in a single forward pass (An et al., 2026; Chen et al., 2026; Liu et al., 2026a), which DDTree then extends into verifiable draft trees (Ringel and Romano, 2026). Concurrent efforts also improve DFlash: Domino (Huang et al., 2026a) introduces a CausalEncoder conceptually similar to our RNN Head, while DFlare (Zhang et al., 2026a) addresses conditioning bottlenecks via layer-wise fusion. System-Aware Scheduling for Speculative Decoding. Beyond drafter architecture, another line of work focuses on determining the optimal number of speculative tokens to generate or verify in each round. To this end, various approaches adapt draft lengths on the fly using confidence heuristics (Du et al., 2024; Li et al., 2024b; Liu et al., 2026c; Mamou et al., 2024; Wen and Feng, 2026), learned acceptance predictors (Huang et al., 2024; Zacks917, 2026), or bandit-style policies (Liu et al., 2026b). Furthermore, recognizing speculative decoding as inherently a systemlevel scheduling problem, recent works optimize overall goodput and latency by adjusting speculation budgets according to real-time system load and request priority (AngelSlim Team, 2026; Hu et al., 2026; Huang et al., 2026b; Li et al., 2026a; Liu et al., 2024c; Miao et al., 2024; Sadhukhan et al., 2025; Wu et al., 2025). Parallel Generation. Models that generate tokens in parallel offer a decoding latency nearly independent of output length, making them an attractive alternative to autoregressive decoding. Non-Autoregressive Transformers (NATs, Gu et al., 2018) pioneered this direction by predicting all positions independently in a single pass. However, this forces the model to average over all 20

在本文中,我们提出了 DSpark,一个旨在克服高并发生产环境中大语言模型推理的结构性和系统级瓶颈的推测解码框架。在算法上,DSpark 引入了一种半自回归生成范式——将计算量大的并行骨干网与轻量级顺序头相结合——以缓解独立并行草稿模型的快速后缀衰减。在系统层面,我们将验证长度选择表述为全局吞吐量最大化问题,采用硬件感知前缀调度器,根据校准的存活概率和实时引擎负载动态调整目标模型的验证预算。广泛的离线评估表明,DSpark 在不同领域上显著优于最先进的自回归和并行基线。此外,其在 DeepSeek-V4 中的实际部署验证了其在生产服务中的实用价值:通过智能管理验证开销,DSpark 在重负载下维持稳健的并发性,持续加速每用户生成速度,并有效将 LLM 服务的帕累托前沿向外移动。

plausible modes, often producing outputs that mix fragments from different valid sequences. Two broad lines of work have emerged to address this limitation. One direction retains the single-pass architecture but changes what the model sees or how it is trained: introducing latent variables as conditioning input to steer all positions toward a consistent output (Gu et al., 2018; Kaiser et al., 2018; Ma et al., 2019), or relaxing the training objective so that the model focuses on producing a single coherent output rather than modeling the full distribution over all valid alternatives (Du et al., 2021; Qian et al., 2021; Shao et al., 2021, 2023). The other direction reintroduces limited sequential dependency through iterative re-prediction (Austin et al., 2021a; Ghazvininejad et al., 2019; Li et al., 2022), block-level autoregression (Arriola et al., 2025; Wang et al., 2018), or structured output layers such as CRF (Sun et al., 2019), CTC (Libovický and Helcl, 2018; Saharia et al., 2020), HMM (Huang et al., 2022b), and PCFG (Gui et al., 2023). Speculative decoding places a further demand that the drafter must provide exact per-token probabilities for the rejection sampling rule. Most techniques above cannot readily provide such probabilities due to iterative refinement, latent marginalization, or global normalization. For instance, in a design closely related to ours, CRF-NAT (Sun et al., 2019) also places a sequential module over parallel hidden states, but its globally normalized partition function prevents exact per-token probability computation. Similarly, when adapting the CTC output layer to parallel speculative decoding, CTC-drafter (Wen et al., 2024) is restricted to greedy verification due to the latent marginalization of alignment paths. DSpark circumvents these limitations by keeping the sequential correction local, so per-token probabilities remain exact softmax evaluations.

结论 Conclusion

In this paper, we present DSpark, a speculative decoding framework designed to overcome the structural and system-level bottlenecks of large language model inference in high-concurrency production environments. Algorithmically, DSpark introduces a semi-autoregressive generation paradigm—coupling a computationally heavy parallel backbone with a lightweight sequential head—to mitigate the rapid suffix decay of independent parallel drafters. At the system level, we formulate verification length selection as a global throughput maximization problem, employing a hardware-aware prefix scheduler that dynamically tailors the target model’s verification budget based on calibrated survival probabilities and real-time engine load. Extensive offline evaluations demonstrate that DSpark substantially outperforms state-of-the-art autoregressive and parallel baselines across diverse domains. Furthermore, its real-world deployment within the DeepSeek-V4 validates its practical value in production serving: by intelligently managing verification overhead, DSpark sustains robust concurrency under heavy load, consistently accelerates per-user generation speeds, and effectively shifts the Pareto frontier of LLM serving outward.

互动版:图/公式 + 针对本篇提问 →