Skip to content
KernelIndex
Search⌘K

Cross-entropy loss

24 eligible runs
cross-entropy

Cross-entropy loss over tokens×vocab logits against integer targets; the benchmark generates fp32 logits.

Fastest reported · usable today
532.4µsmedian · 1.64× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2024-09-03. Reported by source; not independently reproduced.

pip install "liger-kernel==0.2.1"

Fastest by languagetriton · Liger · 532.4 µs · #1python · Transformers · 875.1 µs · 1.64×

Current records

Source-native comparison · GPU NVIDIA A100 · Workload vocab = 4096 · tokens = 16384 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2024-09-03Record history →
Estimated floor 131.7 µs · record 4.04× above itestimate, not evidence ›
DRAM 131.7 µs · bandwidth-bound on A100 80GB
every declared tensor crosses HBM exactly once (2,039 GB/s, A100 80GB datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
532.4µs
1.00×
Reported · BSD-2-Clause · source
2024-09-03stale

1.64× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
TransformersbaselineHugging Face Transformers
875.1µs
1.64×
Reported · Apache-2.0 · source
2024-09-03stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by vocab

1.00 ms10.0 ms40968k16k32k64k128kvocab →532.4 µs810.1 µs1.43 ms2.84 ms6.81 ms15.0 msLiger875.1 µs1.19 ms1.95 ms5.32 ms10.6 ms20.7 msTransformers
LigerTransformersbest per workload · NVIDIA A100 · Liger-Kernel benchmark scripts held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
TransformersHugging Face Transformers
python
875.1µs
1.64×
Reported
Apache-2.0 · source
LigerLiger-Kernel
triton
532.4µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
logitsfloat [tokens, vocab]
targetint64 [tokens]
lossfloat [1]
Axes and behavior
vocabvariable
tokensvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha25614cde8c616cb…
Sources: Liger-Kernel benchmarks (2024-09-03) · BSD-2-Clauselast observed 2024-09-03How records are decidedJSON