NANO-VLLM LAB

Prefill 与 Decode 的两种工作负载

Prefill 一次处理许多 prompt token;Decode 每个活跃序列每轮通常只生成一个新 token。

概念模拟,不执行真实模型

概念时间线

Prefill:并行处理 prompt
Decode:逐轮生成
Prefill token 工作量512
Decode step 数24
每轮活跃序列8