Skip to content
KernelIndex
Search⌘K

GRPO loss, token-level importance sampling

48 eligible runs
grpo

GRPO policy loss over batch×(seq+1)×vocab logits with per-sequence advantages and token-level importance sampling (clip 0.2, no KL reference term).

Fastest reported · usable today
1.81msmedian · 12.13× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-08-05. Reported by source; not independently reproduced.

pip install "liger-kernel==0.6.1"

Fastest by languagetriton · Liger · 1.81 ms · #1python · PyTorch · 22.0 ms · 12.13×

Current records

Source-native comparison · GPU NVIDIA A100 · Workload seq = 1024 · batch = 2 · vocab = 128256 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 4 results · last observed 2025-08-05Record history →
Estimated floor 257.9 µs · record 7.04× above itestimate, not evidence ›
DRAM 257.9 µs · bandwidth-bound on A100 80GB
every declared tensor crosses HBM exactly once (2,039 GB/s, A100 80GB datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
1.81ms
1.00×
Reported · BSD-2-Clause · source
2025-08-05stale

12.13× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
LigerLiger-Kernel
1.82ms
1.01×
Reported · BSD-2-Clause · source
2025-08-05stale

12.07× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
3
PyTorchbaseline
22.0ms
12.13×
Reported · BSD-3-Clause · source
2025-08-05stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →
4
PyTorchbaseline
22.0ms
12.14×
Reported · BSD-3-Clause · source
2025-08-05stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →

Scaling by batch

10.0 ms100.0 ms24816batch →1.81 ms1.85 ms1.89 ms1.97 msLiger22.0 ms41.5 ms81.2 ms160.8 msPyTorch
LigerPyTorchbest per workload · NVIDIA A100 · Liger-Kernel benchmark scripts held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
22.0ms
12.13×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
1.81ms
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
logitsfloat [batch, seq_next, vocab]
completion_idsint64 [batch, seq]
advantagesfloat [batch]
lossfloat [1]
Axes and behavior
seqvariable
batchvariable
vocabvariable
seq_nextseq_next = seq + 1
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2567643fcc6ed3e…
Sources: Liger-Kernel benchmarks (2025-08-05) · BSD-2-Clauselast observed 2025-08-05How records are decidedJSON