model: support speculators-format checkpoints for DSpark (#26275)

* dspark: support speculators-format checkpoints (SpecForge exports)

Speculators-format DSpark drafts (e.g. SpecForge exports for the
Gemma-4-26B-A4B target) differ from the dense DeepSpec checkpoints in
three ways:

- the config nests the backbone hparams under transformer_layer_config
  and gives the extract layers as aux_hidden_state_layer_ids
- the block is the DFlash 1+N fill-in layout: the anchor slot is a bonus
  token, not a prediction slot. Written as dflash.bonus_anchor; such
  drafts build the block and read the mask positions exactly like
  DFlash (n_max drafts from a 1+n_max block), only the Markov/confidence
  sampling comes from DSpark
- the draft output vocab may be reduced (draft_vocab_size < vocab_size)
  with a d2t remap table. The converter expands lm_head/markov_w2 back
  to the full vocab and synthesizes an lm_head bias of -1e9 on the rows
  the draft cannot produce, so the runtime needs no d2t remapping. Such
  drafts ship their own (now optional) token_embd/output tensors instead
  of sharing the target's

Verified against gemma4-26b-a4b-dspark: greedy outputs are byte-identical
with and without the draft; acceptance 0.46, mean draft len 3.7 (n_max 6).

Co-authored-by: desovo7 <942845546@qq.com>
Assisted-by: Claude Fable 5

* dspark: fold the speculators draft class into DSparkModel

One class now covers every DSpark variant. What used to pick the class is
a single flag, because the arch name turns out to be the only thing that
separates the two families: SpecForge also exports a flat schema that
carries no speculators_* fields yet still uses the 1+N bonus-anchor block,
so keying on those fields would silently mis-read its drafts.

Also rename i0 to i_first_pred in the draft read loop and the Markov head,
and give the head a real bonus_anchor bool instead of testing i0 > 0.

Converting the Qwen3-8B DeepSpec draft and both gemma-4 speculators drafts
produces byte-identical GGUFs. The one behaviour change is that the
markov_head_type check now also covers the DeepSpec checkpoints, which
previously skipped it.

Co-authored-by: desovo7 <942845546@qq.com>
Assisted-by: Claude Opus 5

* dspark: address review comments

- rename bonus_anchor to sample_from_anchor (GGUF key and code), matching
  the checkpoint config field; absent key still means anchor-first
- rework the reduced draft vocab to match EAGLE3: d2t is written as I64
  absolute target ids and the logits are scattered at runtime, instead of
  expanding lm_head/markov_w2 and synthesizing an output bias at conversion
- move the t2d skip to modify_tensors, like EAGLE3
- drop _is_specforge: the arch name only picks the sample_from_anchor
  default, embed/lm_head sharing is decided by the draft vocab size
- deduplicate the tok_embd create_tensor left behind by the rebase

Verified with the RedHat gemma-4-31b speculator draft: greedy output is
byte-identical with and without the draft; acceptance 0.26 (n_max 7).

Co-authored-by: desovo7 <942845546@qq.com>
Assisted-by: Claude Fable 5

* dspark: fold the sample_from_anchor read into the block_size block

* dspark: fix flake8 continuation indent

* clean up

* dspark: key the sample_from_anchor default off the export format

  Co-authored-by: desovo7 <942845546@qq.com>
  Assisted-by: Claude Fable

* dspark: drop t2d in filter_tensors

  Co-authored-by: desovo7 <942845546@qq.com>
  Assisted-by: Claude Fable

* dspark: map model.lm_head instead of bypassing the dflash prefix

  Co-authored-by: desovo7 <942845546@qq.com>
  Assisted-by: Claude Fable

---------

Co-authored-by: desovo7 <942845546@qq.com>
Co-authored-by: ruixiang63 <wangruixiang07@outlook.com>
This commit is contained in:
王金旭
2026-08-17 13:51:06 +02:00
committed by GitHub
co-authored by desovo7 ruixiang63
parent 7077abbe14
commit 9cd719af21
8 changed files with 154 additions and 25 deletions
+73 -12
View File
@@ -4,12 +4,13 @@ import json
from typing import Any, Callable, Iterable, TYPE_CHECKING
import numpy as np
import torch
if TYPE_CHECKING:
from torch import Tensor
from .base import ModelBase, TextModel, gguf, logger
from .base import LazyTorchTensor, ModelBase, TextModel, gguf, logger
@ModelBase.register("QWenLMHeadModel")
@@ -708,22 +709,82 @@ class DFlashModel(Qwen3Model):
yield from super().modify_tensors(data_torch, name, bid)
@ModelBase.register("Qwen3DSparkModel")
@ModelBase.register("Qwen3DSparkModel", "DSparkDraftModel", "DSparkSpeculator")
@ModelBase.example("satgeze/Qwen3.6-27B-DSpark")
class DSparkModel(DFlashModel):
# DSpark = DFlash + a semi-autoregressive Markov head
# DSpark = DFlash + a semi-autoregressive Markov head.
model_arch = gguf.MODEL_ARCH.DFLASH
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
# normalize the flat DeepSpec schema to DFlash's nested dflash_config
self.hparams.setdefault("dflash_config", {
k: self.hparams[k] for k in ("target_layer_ids", "mask_token_id") if k in self.hparams
})
def __init__(self, dir_model, *args, **kwargs):
hparams = kwargs.pop("hparams", None)
if hparams is None:
hparams = ModelBase.load_hparams(dir_model, False)
# EAGLE3-style exports use the 1+N bonus-anchor block, DFlash-lineage exports sample from the anchor
self._sample_from_anchor = hparams.get(
"sample_from_anchor",
"transformer_layer_config" not in hparams and "aux_hidden_state_layer_ids" not in hparams)
if "transformer_layer_config" in hparams:
hparams = {**hparams, **hparams["transformer_layer_config"]}
super().__init__(dir_model, *args, hparams=hparams, **kwargs)
# normalize both schemas to DFlash's nested dflash_config
if "aux_hidden_state_layer_ids" in self.hparams:
self.hparams.setdefault("dflash_config", {
"mask_token_id": self.hparams.get("mask_token_id"),
"target_layer_ids": [i - 1 for i in self.hparams["aux_hidden_state_layer_ids"]],
})
else:
self.hparams.setdefault("dflash_config", {
k: self.hparams[k] for k in ("target_layer_ids", "mask_token_id") if k in self.hparams
})
if (markov_head_type := self.hparams.get("markov_head_type", "vanilla")) != "vanilla":
raise ValueError(f"unsupported markov_head_type {markov_head_type!r} (only 'vanilla' is supported)")
n_vocab = self.hparams["vocab_size"]
self._n_vocab_draft = self.hparams.get("draft_vocab_size") or n_vocab
if self._n_vocab_draft > n_vocab:
raise ValueError(f"draft_vocab_size {self._n_vocab_draft} exceeds vocab_size {n_vocab}")
self._d2t: Tensor | None = None
def set_gguf_parameters(self):
super().set_gguf_parameters()
self.gguf_writer.add_sample_from_anchor(self._sample_from_anchor)
@classmethod
def filter_tensors(cls, item: tuple[str, Callable[[], Tensor]]) -> tuple[str, Callable[[], Tensor]] | None:
name, gen = item
if name.endswith(("embed_tokens.weight", "lm_head.weight")):
if item[0] == "t2d": # not used at runtime
return None
return super().filter_tensors((name, gen))
return super().filter_tensors(item)
def modify_tensors(self, data_torch: Tensor, name: str, bid: int | None) -> Iterable[tuple[str, Tensor]]:
if name == "model.d2t":
self._d2t = data_torch
return
if self._n_vocab_draft == self.hparams["vocab_size"] and name.endswith(("embed_tokens.weight", "lm_head.weight")):
return
yield from super().modify_tensors(data_torch, name, bid)
def prepare_tensors(self):
super().prepare_tensors()
n_vocab = self.hparams["vocab_size"]
if self._n_vocab_draft < n_vocab and self._d2t is None:
raise ValueError(f"draft_vocab_size {self._n_vocab_draft} < vocab_size {n_vocab} but no d2t table found")
# write d2t as absolute target token ids
if self._d2t is not None:
data = LazyTorchTensor.to_eager(self._d2t).to(torch.int64).cpu().numpy().reshape(-1)
if data.size != self._n_vocab_draft:
raise ValueError(f"d2t size {data.size} does not match draft_vocab_size {self._n_vocab_draft}")
data = data + np.arange(data.size, dtype=np.int64)
if np.any((data < 0) | (data >= n_vocab)):
raise ValueError(f"d2t target ids out of range for target vocab size {n_vocab}")
if np.unique(data).size != data.size:
raise ValueError("d2t contains duplicate target ids")
logger.info(f"{'d2t,':<30} --> I64, shape = {{{data.size}}}")
self.gguf_writer.add_tensor("d2t", data, raw_dtype=gguf.GGMLQuantizationType.I64)