Skip to content
KernelIndex
Search⌘K

Total variation distance loss

36 eligible runs
tvd

Total variation distance between two tokens×vocab probability distributions (batchmean reduction); the benchmark generates fp32 inputs.

Fastest reported · usable today
275.7µsmedian · 2.64× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2026-03-03. Reported by source; not independently reproduced.

pip install "liger-kernel==0.7.0"

Fastest by languagetriton · Liger · 275.7 µs · #1python · PyTorch · 728.8 µs · 2.64×

Current records

Source-native comparison · GPU NVIDIA H100 · Workload vocab = 4096 · tokens = 16384 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2026-03-03Record history →
Estimated floor 160.3 µs · record 1.72× above itestimate, not evidence ›
DRAM 160.3 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
275.7µs
1.00×
Reported · BSD-2-Clause · source
2026-03-03stale

2.64× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
PyTorchbaseline
728.8µs
2.64×
Reported · BSD-3-Clause · source
2026-03-03stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →

Scaling by vocab

1.00 ms10.0 ms40968k16k32k64k128kvocab →275.7 µs533.9 µs1.05 ms2.10 ms4.22 ms8.50 msLiger728.8 µs1.43 ms2.81 ms5.60 ms11.2 ms22.3 msPyTorch
LigerPyTorchbest per workload · NVIDIA H100 · Liger-Kernel benchmark scripts held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
728.8µs
2.64×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
275.7µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
pfloat [tokens, vocab]
qfloat [tokens, vocab]
lossfloat [1]
Axes and behavior
vocabvariable
tokensvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256ae6b46f63ebb…
Sources: Liger-Kernel benchmarks (2026-03-03) · BSD-2-Clauselast observed 2026-03-03How records are decidedJSON