ReLU squared
48 eligible runs
relu-squared
Elementwise relu(x)^2 activation over a rows×hidden tensor; the benchmark sweeps the hidden size.
Fastest reported · usable today
7.33µsmedian · 1.18× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2026-03-27. Reported by source; not independently reproduced.
pip install "liger-kernel==0.7.0"Fastest by languagetriton · Liger · 7.33 µs · #1python · PyTorch · 8.67 µs · 1.18×
Current records
Workloadrows = 4096 · hidden = 4096 · bf16rows = 4096 · hidden = 8192 · bf16rows = 4096 · hidden = 2048 · bf16rows = 4096 · hidden = 16384 · bf16rows = 4096 · hidden = 128 · bf16rows = 4096 · hidden = 256 · bf16rows = 4096 · hidden = 512 · bf16rows = 4096 · hidden = 1024 · bf16
HardwareBest knownImplementationRuns
NVIDIA H100 · env 17.33 µsLiger2NVIDIA H100 · env 27.94 µsLiger2NVIDIA H100 · env 325.9 µsPyTorch2Not measured on B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA H100 · Workload rows = 4096 · hidden = 128 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2026-03-27Record history →
Estimated floor 313 ns · record 23.41× above itestimate, not evidence ›
DRAM 313 ns · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel7.33µs1.00×Reported · BSD-2-Clause · source2026-03-27stale
1.18× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2PyTorchbaseline8.67µs1.18×Reported · BSD-3-Clause · source2026-03-27stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by hidden
Implementations
Semantics
Inputs and outputs
xfloat [rows, hidden]
yfloat [rows, hidden]
Axes and behavior
rowsvariable
hiddenvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256ddb7fe09e7cc…
Sources: Liger-Kernel benchmarks (2026-03-27) · BSD-2-Clauselast observed 2026-03-27How records are decidedJSON