mlx-dspark
==========

This project is an independent MLX/Apple-Silicon port of the *inference path* of
DSpark, a speculative-decoding drafter open-sourced by DeepSeek as part of the
DeepSpec codebase:

  - DeepSpec: https://github.com/deepseek-ai/DeepSpec  (MIT License)

It loads the published DSpark drafter checkpoints (also released by DeepSeek under
their respective licenses):

  - deepseek-ai/dspark_gemma4_12b_block7
  - deepseek-ai/dspark_qwen3_4b_block7

The DSpark drafter architecture, training, and checkpoints are the work of DeepSeek
and the DeepSpec authors. This repository reimplements only the DSpark forward/
verification path for MLX and contains no DeepSpec source code.

---

DFlash (z-lab)
--------------

mlx-dspark also runs z-lab's original DFlash drafter (block diffusion for speculative
decoding):

  - DFlash: https://github.com/z-lab/dflash  (MIT License)
  - Paper:  Chen et al., "DFlash: Block Diffusion for Flash Speculative Decoding",
            arXiv:2602.06036

The file `src/mlx_dspark/dflash_model.py` contains the DFlash drafter *model* classes
(DFlashConfig, DFlashAttention, DFlashDecoderLayer, DFlashDraftModel) vendored verbatim
from z-lab/dflash (`dflash/model_mlx.py`) under the MIT License; copyright (c) z-lab.
The MIT permission notice is reproduced in that file's header. The generation/
verification loop around it (`generate.dflash_generate`) is mlx-dspark's own. DFlash
checkpoints (e.g. `z-lab/gemma4-12B-it-DFlash`) are downloaded at runtime from z-lab.

The same file's **DFlash 2** components (DFlashGroupedConv, CandidateSelector and the
selector drafting path) are an MLX port of the DFlash 2 reference implementation
contributed to SGLang (sgl-project/sglang, `srt/models/dflash.py` and
`srt/speculative/dflash_worker_v2.py`, Apache License 2.0). DFlash 2 is by Inco AI
(https://inco.ai/blog/dflash2/); its checkpoints (e.g. `incoai/Qwen3.8-27B-DFlash2`,
Apache-2.0) are downloaded at runtime from their publisher.

---

Small-M quantized-matmul kernel (credits; no third-party code included)
-----------------------------------------------------------------------

`src/mlx_dspark/skinny_qmm.py` is mlx-dspark's own implementation. Its design builds on
ideas from two MIT-licensed projects (their code is not included):

  - avlp12's mlx-lm fork (`mlx_lm/fast_qmm.py`, the `qmm_mma4` kernel; shipped here
    v0.12.0-v0.19.0): a simdgroup-matrix kernel for the few-row qmm dead zone
    (https://github.com/avlp12/mlx-lm, context: https://github.com/ml-explore/mlx/issues/4265)
  - TensorFold (`simd_qmm`, https://github.com/ashhart/TensorFold): weights as the
    matrix operand for skinny batches, with quantized fields masked in place against
    pre-scaled inputs (a technique also long used by llama.cpp's Metal kernels)

---

Nanbeige model module (Apple / mlx-lm)
--------------------------------------

The file `src/mlx_dspark/nanbeige_lm.py` contains the Nanbeige4.2 model classes
(ModelArgs, Attention, MLP, TransformerBlock, NanbeigeModel, Model) vendored from the
mlx-lm main branch (`mlx_lm/models/nanbeige.py`, MIT License; Copyright © 2026 Apple
Inc.), which no mlx-lm release ships yet:

  - https://github.com/ml-explore/mlx-lm  (MIT License)

The runtime registration shim around it is mlx-dspark's own and self-retires once an
installed mlx-lm provides the module itself.

---

Target models (Gemma-4, Qwen3) are downloaded at runtime from their respective
publishers and are subject to their own licenses (e.g. the Gemma Terms of Use and the
Qwen license).

No model weights are bundled with this package.
