Chunked distillation cosine loss
16 eligible runs
cosine-similarity-loss
Distillation loss combining hard cross-entropy and a soft cosine-similarity term between a hidden/2 student and a hidden teacher, each projected through its own LM head (equal hard/soft weights).
Fastest reported · usable today
13.8msmedian · 1.19× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2025-06-27. Reported by source; not independently reproduced.
pip install "liger-kernel==0.5.10"Fastest by languagetriton · Liger · 13.8 ms · #1python · PyTorch · 16.5 ms · 1.19×
Current records
Workloadvocab = 128256 · hidden = 4096 · tokens = 1024 · bf16vocab = 128256 · hidden = 4096 · tokens = 2048 · bf16vocab = 128256 · hidden = 4096 · tokens = 4096 · bf16vocab = 128256 · hidden = 4096 · tokens = 8192 · bf16
HardwareBest knownImplementationRuns
NVIDIA A100 · env 113.8 msLiger2NVIDIA A100 · env 214.7 msLiger2Not measured on H100, B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA A100 · Workload vocab = 128256 · hidden = 4096 · tokens = 1024 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-06-27Record history →
Estimated floor 779.1 µs · record 17.75× above itestimate, not evidence ›
DRAM 779.1 µs · bandwidth-bound on A100 80GB
every declared tensor crosses HBM exactly once (2,039 GB/s, A100 80GB datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel13.8ms1.00×Reported · BSD-2-Clause · source2025-06-27stale
1.19× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2PyTorchbaseline16.5ms1.19×Reported · BSD-3-Clause · source2025-06-27stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by tokens
Implementations
Semantics
Inputs and outputs
student_inputfloat [tokens, student_hidden]
student_weightfloat [vocab, student_hidden]
teacher_inputfloat [tokens, hidden]
teacher_weightfloat [vocab, hidden]
targetint64 [tokens]
lossfloat [1]
Axes and behavior
vocabvariable
hiddenvariable
tokensvariable
student_hiddenstudent_hidden = hidden // 2
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2563afad1b6dfc5…
Sources: Liger-Kernel benchmarks (2025-06-27) · BSD-2-Clauselast observed 2025-06-27How records are decidedJSON