Model catalog // model

Qwen3.8-27B-DFlash2

496B125F-7AEE-41FF-851B-DE96CB42D7E5 EIM Only

parameters
4.2B
context
262K
license
Unknown
architecture
qwen3
// description

About Qwen3.8-27B-DFlash2

---
license: apache-2.0
libraryname: transformers
pipeline
tag: text-generation
basemodel:

  • Qwen/Qwen3.8-27B

inference: false
tags:
  • dflash2
  • speculative-decoding
  • block-diffusion
  • draft-model
  • sglang
  • vllm

---

Qwen3.8-27B-DFlash2

Blog | GitHub

This repository contains the DFlash 2 draft model for
Qwen/Qwen3.8-27B.
It is not a standalone language model: it runs inside a speculative
decoding server and drafts tokens for the target model to verify. The checkpoint is also
mirrored at z-lab/Qwen3.8-27B-DFlash2.

DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts
a whole block of tokens in a single pass and keeps the top candidates at
every position. A lightweight selector then traces one coherent path through them.
Two-tap dynamic convolutions in the backbone keep the draft from decaying
toward the end of the block. Decoding is lossless: greedy output
matches the target model exactly, and sampling preserves its distribution.

<div align="center">
<img src="assets/dflash2-figure.png" alt="DFlash 2: parallel block drafting with a candidate path selector" width="100%">
</div>

Quick Start

Serve with SGLang:

pip install "sglang[all] @ git+https://github.com/sgl-project/sglang.git#subdirectory=python"

python -m sglang.launch_server \
  --model-path Qwen/Qwen3.8-27B \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 \
  --speculative-num-draft-tokens 8

Or with vLLM:

pip install -U "vllm @ git+https://github.com/vllm-project/vllm.git@refs/pull/52816/head"

vllm serve Qwen/Qwen3.8-27B \
  --speculative-config '{
    "method": "dflash",
    "model": "incoai/Qwen3.8-27B-DFlash2",
    "num_speculative_tokens": 7
  }'

See the blog post for other engines and more details.

Evaluation

  • Runtime: SGLang on one NVIDIA H200, with FlashAttention 3 for target and draft attention
  • Speculation block size: 8 (7 draft tokens per verification step)
  • Sampling: Qwen3.8's officially recommended parameters (temperature 1.0, top-p 0.95, top-k 20), with xhigh reasoning effort
  • Maximum new tokens: 4096
  • Prompts: benchmark formatting from z-lab/dflash

We compare autoregressive decoding, Qwen3.8's built-in seven-token MTP,
a community DSpark drafter
(RadixArk/Qwen3.8-27B-DSpark),
and DFlash 2. All speculative methods propose seven draft tokens per
verification step.

Acceptance Length

Acceptance length is the per-request mean of completion tokens divided by verification steps.
Higher is better.

TaskMTPDSparkDFlash 2
GSM8K5.024.365.46
MATH-5004.723.925.28
HumanEval3.913.304.39
MBPP3.993.514.79
MT-Bench3.743.014.10

Throughput

Throughput is total output tokens divided by end-to-end wall time.
Each cell shows output tok/s (speedup vs. autoregressive).

Concurrency 1

TaskAutoregressiveMTPDSparkDFlash 2
GSM8K68.9178.5 (2.59×)185.3 (2.69×)236.1 (3.43×)
MATH-50069.0172.8 (2.51×)174.5 (2.53×)230.7 (3.34×)
HumanEval69.0151.9 (2.20×)159.9 (2.32×)214.6 (3.11×)
MBPP69.0153.1 (2.22×)163.3 (2.37×)226.9 (3.29×)
MT-Bench68.9134.9 (1.96×)137.6 (2.00×)184.0 (2.67×)

Concurrency 8

TaskAutoregressiveMTPDSparkDFlash 2
GSM8K467.21,022.1 (2.19×)1,040.8 (2.23×)1,328.7 (2.84×)
MATH-500480.01,023.5 (2.13×)1,025.8 (2.14×)1,368.3 (2.85×)
HumanEval483.4934.2 (1.93×)956.5 (1.98×)1,291.5 (2.67×)
MBPP478.0938.1 (1.96×)974.1 (2.04×)1,328.0 (2.78×)
MT-Bench480.5835.2 (1.74×)802.3 (1.67×)1,090.2 (2.27×)

Concurrency 32

TaskAutoregressiveMTPDSparkDFlash 2
GSM8K1,329.81,381.1 (1.04×)1,506.5 (1.13×)1,922.5 (1.45×)
MATH-5001,505.81,415.6 (0.94×)1,429.0 (0.95×)1,951.8 (1.30×)
HumanEval1,546.51,296.8 (0.84×)1,330.1 (0.86×)1,799.0 (1.16×)
MBPP1,507.71,314.9 (0.87×)1,361.3 (0.90×)1,886.8 (1.25×)
MT-Bench1,507.41,159.7 (0.77×)1,115.5 (0.74×)1,525.3 (1.01×)

Citation

If you find DFlash 2 useful, please cite:

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}

Please also cite the original DFlash paper:

@inproceedings{chen2026dflash,
  title     = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author    = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  booktitle = {International Conference on Machine Learning (ICML)},
  year      = {2026}
}

// specs

Specifications

Core capabilities and constraints for this model.

Parameters 4.2B
Context 262K
Architecture qwen3
License Unknown
// usage

Use this model

Deploy and call it with an OpenAI-compatible request.

Shell
# Create an endpoint for this model, then provision it. # --tier-id is the service tier to run on; pick one at /endpoints/new. erectl endpoints create \ --name "My endpoint" \ --slug my-endpoint \ --model-id 496B125F-7AEE-41FF-851B-DE96CB42D7E5 \ --tier-id <tier-id> # provision takes the endpoint UUID returned by create erectl endpoints provision <endpoint-id> # Call it. Requests are routed by /<project-id>/<endpoint-slug>. curl https://api.erebine.ai/proj_ABC123/my-endpoint/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $EREBINE_API_KEY" \ -d '{ "model": "Qwen3.8-27B-DFlash2", "messages": [{"role": "user", "content": "Hello, world!"}], "stream": true }'
// tags

Tags

Workload types and capability tags.

chat reasoning

Need guaranteed availability?

Deploy this model on dedicated EIM nodes for production workloads.