NVIDIA LPU supports 3 types of disaggregated inferencing:
1. Rubin Prefill + LPU Decode for the fastest interactivity2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curveFor low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.


