Skip to content
KernelIndex
Search⌘K

Chunked distillation JSD loss

16 eligible runs
jsd

Distillation loss combining hard cross-entropy and soft JSD between a hidden/2 student and a hidden teacher, each projected through its own LM head (equal hard/soft weights).

Fastest reported · usable today
7.74msmedian · 1.41× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2024-12-03. Reported by source; not independently reproduced.

pip install "liger-kernel==0.4.2"

Fastest by languagetriton · Liger · 7.74 ms · #1python · PyTorch · 10.9 ms · 1.41×

Current records

Source-native comparison · GPU NVIDIA H100 · Workload vocab = 128256 · hidden = 4096 · tokens = 1024 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2024-12-03Record history →
Estimated floor 474.2 µs · record 16.31× above itestimate, not evidence ›
DRAM 474.2 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
7.74ms
1.00×
Reported · BSD-2-Clause · source
2024-12-03stale

1.41× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
PyTorchbaseline
10.9ms
1.41×
Reported · BSD-3-Clause · source
2024-12-03stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →

Scaling by tokens

20.0 ms40.0 ms60.0 ms80.0 ms1024204840968ktokens →7.74 ms15.2 ms30.2 ms60.2 msLiger10.9 ms21.5 ms43.0 ms85.4 msPyTorch
LigerPyTorchbest per workload · NVIDIA H100 · Liger-Kernel benchmark scripts held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
10.9ms
1.41×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
7.74ms
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
student_inputfloat [tokens, student_hidden]
student_weightfloat [vocab, student_hidden]
teacher_inputfloat [tokens, hidden]
teacher_weightfloat [vocab, hidden]
targetint64 [tokens]
lossfloat [1]
Axes and behavior
vocabvariable
hiddenvariable
tokensvariable
student_hiddenstudent_hidden = hidden // 2
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2564ce30dbf7da7…
Sources: Liger-Kernel benchmarks (2024-12-03) · BSD-2-Clauselast observed 2024-12-03How records are decidedJSON