Skip to content
KernelIndex
Search⌘K

Fused linear + JSD

16 eligible runs
jsd

Student and teacher tokens×hidden activations projected through separate vocab×hidden heads fused with generalized JSD (beta 0.5, batchmean) between the resulting log-distributions.

Best usable
110.0msmedian · 11.51× vs fastest
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2024-10-09. Reported by source; not independently reproduced.

pip install "liger-kernel==0.3.1"
Source baseline · unbeaten · not usable as-is
9.56msmedian
PyTorchBSD-3-Clause · python

Reported evidence · last observed 2024-10-09. The source's designated baseline implementation. Reported by source; not independently reproduced.

Fastest by languagepython · PyTorch · 9.56 ms · #1triton · Liger · 110.0 ms · 11.51×

Current records

Source-native comparison · GPU NVIDIA H100 · Workload vocab = 128256 · hidden = 4096 · tokens = 1024 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2024-10-09Record history →
Estimated floor 632.3 µs · record 15.12× above itestimate, not evidence ›
DRAM 632.3 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
PyTorchbaseline
9.56ms
1.00×
Reported · BSD-3-Clause · source
2024-10-09stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →
2
LigerLiger-Kernel
110.0ms
11.51×
Reported · BSD-2-Clause · source
2024-10-09stale

11.51× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →

Scaling by tokens

50.0 ms100.0 ms150.0 ms1024204840968ktokens →9.56 ms18.7 ms37.8 ms75.2 msPyTorch110.0 ms124.1 ms143.2 ms180.9 msLiger
PyTorchLigerbest per workload · NVIDIA H100 · Liger-Kernel benchmark scripts held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
9.56ms
1.00×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
110.0ms
11.51×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
student_inputfloat [tokens, hidden]
student_weightfloat [vocab, hidden]
teacher_inputfloat [tokens, hidden]
teacher_weightfloat [vocab, hidden]
lossfloat [1]
Axes and behavior
vocabvariable
hiddenvariable
tokensvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha25671138de338d4…
Sources: Liger-Kernel benchmarks (2024-10-09) · BSD-2-Clauselast observed 2024-10-09How records are decidedJSON