Spec Decoding (DSpark) finally works with Pipeline Parallelism in vLLM!! 🔥 A lot of the GPU poor/proletarian have been asking for this feature for a while now, but it took the recent Kimi K3's massive 2.8T parameters affecting the GPU middle class (B200) for it to be implemented!
Before this change, GPU poors that needed to use pipeline parallelism, as their weights didn't fully fit on 1 server, would not be able to take advantage of DSpark.Shoutout to yongqinwang-cmd & Inferact for implementing it in vLLM!

