
Parallel All the Way Down: Beyond Single-Token Generation with Speculative Decoding
·7 min read
Speculators and vLLM now support P-EAGLE, DFlash, and DSpark — three parallel drafting algorithms that move beyond sequential token generation to deliver faster, simpler, and more scalable speculative decoding for LLM serving.