RMSNorm
36 eligible runs
rmsnorm
Root-mean-square normalization of a rows×hidden activation with a learned hidden-size weight.
Fastest reported · usable today
13.6µsmedian · 5.75× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2024-09-03. Reported by source; not independently reproduced.
pip install "liger-kernel==0.2.1"Fastest by languagetriton · Liger · 13.6 µs · #1python · Transformers · 78.2 µs · 5.75×
Current records
Workloadrows = 2048 · hidden = 1024 · bf16rows = 2048 · hidden = 2048 · bf16rows = 2048 · hidden = 4096 · bf16rows = 2048 · hidden = 8192 · bf16rows = 2048 · hidden = 16384 · bf16rows = 2048 · hidden = 32768 · bf16
HardwareBest knownImplementationRuns
NVIDIA A100 · env 113.6 µsLiger2NVIDIA A100 · env 293.7 µsLiger2NVIDIA A100 · env 3286.0 µsLiger2Not measured on H100, B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA A100 · Workload rows = 2048 · hidden = 1024 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2024-09-03Record history →
Estimated floor 2.06 µs · record 6.61× above itestimate, not evidence ›
DRAM 2.06 µs · bandwidth-bound on A100 80GB
every declared tensor crosses HBM exactly once (2,039 GB/s, A100 80GB datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel13.6µs1.00×Reported · BSD-2-Clause · source2024-09-03stale
5.75× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
278.2µs5.75×Reported · Apache-2.0 · source2024-09-03stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredApache-2.0no install recipeView source →Run detail →
Scaling by hidden
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
xfloat [rows, hidden]
weightfloat [hidden]
yfloat [rows, hidden]
Axes and behavior
rowsvariable
hiddenvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2568ba8d8c445e8…
Sources: Liger-Kernel benchmarks (2024-09-03) · BSD-2-Clauselast observed 2024-09-03How records are decidedJSON