KL-divergence loss
24 eligible runs
kl-divergence
KL divergence between tokens×vocab log-probabilities and target probabilities (batchmean reduction); the benchmark generates fp32 inputs.
Fastest reported · usable today
306.4µsmedian · 4.34× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2024-09-04. Reported by source; not independently reproduced.
pip install "liger-kernel==0.2.1"Fastest by languagetriton · Liger · 306.4 µs · #1python · PyTorch · 1.33 ms · 4.34×
Current records
Workloadvocab = 4096 · tokens = 16384 · fp32vocab = 8192 · tokens = 16384 · fp32vocab = 16384 · tokens = 16384 · fp32vocab = 32768 · tokens = 16384 · fp32vocab = 131072 · tokens = 16384 · fp32vocab = 65536 · tokens = 16384 · fp32
HardwareBest knownImplementationRuns
NVIDIA H100 · env 1306.4 µsLiger2NVIDIA H100 · env 22.04 msLiger2Not measured on B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA H100 · Workload vocab = 4096 · tokens = 16384 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2024-09-04Record history →
Estimated floor 160.3 µs · record 1.91× above itestimate, not evidence ›
DRAM 160.3 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel306.4µs1.00×Reported · BSD-2-Clause · source2024-09-04stale
4.34× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2PyTorchbaseline1.33ms4.34×Reported · BSD-3-Clause · source2024-09-04stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by vocab
Implementations
Semantics
Inputs and outputs
log_probsfloat [tokens, vocab]
target_probsfloat [tokens, vocab]
lossfloat [1]
Axes and behavior
vocabvariable
tokensvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2566304060bf19a…
Sources: Liger-Kernel benchmarks (2024-09-04) · BSD-2-Clauselast observed 2024-09-04How records are decidedJSON