GRPO loss, token-level importance sampling
GRPO policy loss over batch×(seq+1)×vocab logits with per-sequence advantages and token-level importance sampling (clip 0.2, no KL reference term).
Reported evidence · last observed 2025-08-05. Reported by source; not independently reproduced.
pip install "liger-kernel==0.6.1"Fastest by languagetriton · Liger · 1.81 ms · #1python · PyTorch · 22.0 ms · 12.13×
Current records
Not measured on H100, B200 for this workload. Challenges →
Estimated floor 257.9 µs · record 7.04× above itestimate, not evidence ›
1LigerLiger-Kernel1.81ms1.00×Reported · BSD-2-Clause · source2025-08-05stale
12.13× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
2LigerLiger-Kernel1.82ms1.01×Reported · BSD-2-Clause · source2025-08-05stale
12.07× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
3PyTorchbaseline22.0ms12.13×Reported · BSD-3-Clause · source2025-08-05stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
4PyTorchbaseline22.0ms12.14×Reported · BSD-3-Clause · source2025-08-05stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.