Vocab-parallel cross-entropy, Megatron layout
54 eligible runs
cross-entropy
Vocab-parallel cross-entropy over seq×batch×vocab bf16 logits against integer targets, returning per-token fp32 losses. Only TP=1 rows import; sharded-vocab rows would change the local logits shape.
Fastest reported · usable today
138.6µsmedian
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2026-06-15. Reported by source; not independently reproduced.
pip install "liger-kernel==0.8.0"Fastest by languagetriton · Liger · 138.6 µs · #1python · Megatron fused · 312.2 µs · 2.25×
Current records
Workloadseq = 2048 · batch = 4 · vocab = 4096 · bf16seq = 2048 · batch = 4 · vocab = 65536 · bf16seq = 2048 · batch = 4 · vocab = 131072 · bf16seq = 2048 · batch = 4 · vocab = 8192 · bf16seq = 2048 · batch = 4 · vocab = 16384 · bf16seq = 2048 · batch = 4 · vocab = 32768 · bf16
HardwareBest knownImplementationRuns
NVIDIA H100 · env 1138.6 µsLiger3NVIDIA H100 · env 2349.4 µsLiger3NVIDIA H100 · env 3349.8 µsLiger3Not measured on B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA H100 · Workload seq = 2048 · batch = 4 · vocab = 4096 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 3 results · last observed 2026-06-15Record history →
Estimated floor 20.1 µs · record 6.91× above itestimate, not evidence ›
DRAM 20.1 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel138.6µs1.00×Reported · BSD-2-Clause · source2026-06-15
Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2Megatron fusedNVIDIA Megatron-LM312.2µs2.25×Reported · BSD-3-Clause · source2026-06-15
Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
3Megatron fusedNVIDIA Megatron-LM558.5µs4.03×Reported · BSD-3-Clause · source2026-06-15
Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by vocab
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
logitsfloat [seq, batch, vocab]
targetint64 [seq, batch]
lossfp32 [seq, batch]
Axes and behavior
seqvariable
batchvariable
vocabvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2569efef198c06e…
Sources: Liger-Kernel benchmarks (2026-06-15) · BSD-2-Clauselast observed 2026-06-15How records are decidedJSON