Skip to content
KernelIndex
Search⌘K

Fused add RMSNorm h5120

10 eligible runs
rmsnorm

Fused Add + RMSNorm with hidden_size=5120 for Qwen3 14B. Epsilon is fixed at 1e-6.

Source baseline · unbeaten
3.30µsmean
flashinfer / wrapper7f6051FlashInfer-Bench baselines · Apache-2.0 · python

Reported evidence · last observed 2026-03-26. The source's designated baseline implementation. Reported by source; not independently reproduced.

Current records

Source-native comparison · GPU NVIDIA B200 · Workload batch_size = 32 · bf16 · CUDA 12.8 · Framework pytorch 2.9.1+cu128 · Protocol flashinfer-bench · mean · 2 results · last observed 2026-03-26Record history →
Estimated floor 83 ns · record 39.62× above itestimate, not evidence ›
DRAM 83 ns · bandwidth-bound on B200
every declared tensor crosses HBM exactly once (8,000 GB/s, B200 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
flashinfer / wrapper7f6051baselineFlashInfer-Bench baselines
3.30µs
1.00×
Reported · Apache-2.0 · source
2026-03-26stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
2
flashinfer / wrapper7f6051baselineFlashInfer-Bench baselines
3.42µs
1.04×
Reported · Apache-2.0 · source
2026-03-24stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by batch_size

5.00 µs10.0 µs1632371174batch_size →3.58 µs3.33 µs3.30 µs3.42 µs12.7 µsflashinfer / wrapper 7f6051
flashinfer / wrapper 7f6051best per workload · NVIDIA B200 · CUDA 12.8 · pytorch 2.9.1+cu128 · flashinfer-bench held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
flashinfer / wrapper7f6051FlashInfer-Bench baselines
python
3.30µs
1.00×
Reported
Apache-2.0 · source

Semantics

Inputs and outputs
hidden_statesbf16 [batch_size, hidden_size]
residualbf16 [batch_size, hidden_size]
weightbf16 [hidden_size]
outputbf16 [batch_size, hidden_size]
Axes and behavior
batch_sizevariable
hidden_sizeconstant = 5120
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasfused_add_rmsnorm_h5120modelqwen3-14bsha2566cd4d96033b7…
Sources: FlashInfer-Bench (2026-03-26) · Apache-2.0last observed 2026-03-26How records are decidedJSON