veneta

News · 2026-10-07

A Korean draft vocabulary for MTP speculative decoding, measured and published

Qwen3.8-Flash-Next decodes Korean at 16.9 to 26.1 tokens per second (+54%) with a Korean draft vocabulary for MTP speculative decoding, measured on one DGX Spark and published under Apache-2.0.

VENETA Inc. · Open source

Local models with an MTP head draft several tokens at once by restricting the draft head's output to a frequency-ranked vocabulary subset; the target model still verifies over the full vocabulary, so the output is unchanged. The subsets the community ships today are ranked on English and code, so Korean drafts are rejected and decoding slows down for Korean answers.

What we measured

On Qwen3.8-Flash-Next (nvidia/Qwen3.8-Flash-Next-NVFP4, one DGX Spark, median of three runs), a Korean draft vocabulary built the same way as six other languages already published by the community: decode 16.9 → 26.1 tokens/sec (+54%), accepted tokens per draft 1.33 → 2.16. Quality held: zero replacement characters across every generation, no broken Korean script, no change on our 12-case memory-ability regression suite.

What is published, and what is not

The vocabulary file, the build script and every raw run file are public under Apache-2.0: GitHub, Hugging Face. Only Qwen3.8-Flash-Next is measured today; other model families and a generic patch for separate-drafter architectures are in progress, not yet published.

The note is on the project site: MTP speculative decoding.

All news