Fused add + RMSNorm
54 eligible runs
rmsnorm
Residual add followed by RMSNorm, returning the normalized activation and the post-add residual.
Fastest reported · usable today
20.5µsmedian · 2.22× faster than baseline
Liger fusedLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2026-04-07. Reported by source; not independently reproduced.
pip install "liger-kernel==0.7.0"Fastest by languagetriton · Liger fused · 20.5 µs · #1python · Transformers · 45.5 µs · 2.22×
Current records
Workloadrows = 2048 · hidden = 1024 · fp32rows = 2048 · hidden = 8192 · fp32rows = 2048 · hidden = 16384 · fp32rows = 2048 · hidden = 32768 · fp32rows = 2048 · hidden = 2048 · fp32rows = 2048 · hidden = 4096 · fp32
HardwareBest knownImplementationRuns
NVIDIA H100 · env 120.5 µsLiger fused3NVIDIA H100 · env 2113.6 µsTransformers3NVIDIA H100 · env 3229.6 µsTransformers3Not measured on B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA H100 · Workload rows = 2048 · hidden = 1024 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 3 results · last observed 2026-04-07Record history →
Estimated floor 5.01 µs · record 4.1× above itestimate, not evidence ›
DRAM 5.01 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1Liger fusedLiger-Kernel20.5µs1.00×Reported · BSD-2-Clause · source2026-04-07
2.22× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
245.5µs2.22×Reported · Apache-2.0 · source2026-04-07
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredApache-2.0no install recipeView source →Run detail →
3Liger fusedLiger-Kernel46.1µs2.24×Reported · BSD-2-Clause · source2026-04-07
1.01× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
Scaling by hidden
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
xfloat [rows, hidden]
residualfloat [rows, hidden]
weightfloat [hidden]
yfloat [rows, hidden]
residual_outfloat [rows, hidden]
Axes and behavior
rowsvariable
hiddenvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256770a9f88bcb4…
Sources: Liger-Kernel benchmarks (2026-04-07) · BSD-2-Clauselast observed 2026-04-07How records are decidedJSON