@

NVIDIA LPU supports 3 types of disaggregated inferencing:

1. Rubin Prefill + LPU Decode for the fastest interactivity

2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve

3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve

For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.

图片 1,共 1 张

The copyright of this article belongs to the original author/organization.

The views expressed herein are solely those of the author and do not reflect the stance of the platform. The content is intended for investment reference purposes only and shall not be considered as investment advice. Please contact us if you have any questions or suggestions regarding the content services provided by the platform.