Skip to content
KernelIndex
Search⌘K

SwiGLU MLP

68 eligible runs
swiglu

Llama-style SwiGLU MLP: silu(x·Wgate)·(x·Wup) projected back down, over batch×seq×hidden activations.

Fastest reported · usable today
659.5µsmedian · 1.02× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-11-11. Reported by source; not independently reproduced.

pip install "liger-kernel==0.6.3"

Fastest by languagetriton · Liger · 659.5 µs · #1python · Transformers · 674.8 µs · 1.02×

Current records

Source-native comparison · GPU NVIDIA GeForce RTX 4090 · Workload seq = 1024 · batch = 2 · hidden = 2048 · intermediate = 4096 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 4 results · last observed 2025-11-11Record history →
Estimated floor 58.3 µs · record 11.32× above itestimate, not evidence ›
DRAM 58.3 µs · bandwidth-bound on RTX 4090
every declared tensor crosses HBM exactly once (1,008 GB/s, RTX 4090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
659.5µs
1.00×
Reported · BSD-2-Clause · source
2025-11-11stale

1.02× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
TransformersbaselineHugging Face Transformers
674.8µs
1.02×
Reported · Apache-2.0 · source
2025-11-11stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
3
LigerLiger-Kernel
739.5µs
1.12×
Reported · BSD-2-Clause · source
2025-11-11stale

1.10× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
4
TransformersbaselineHugging Face Transformers
745.5µs
1.13×
Reported · Apache-2.0 · source
2025-11-11stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by seq

5.00 ms10.0 ms1024204840968k16kseq →659.5 µs1.35 ms2.72 ms5.34 ms10.9 msLiger674.8 µs1.41 ms2.83 ms5.66 ms11.3 msTransformers
LigerTransformersbest per workload · NVIDIA GeForce RTX 4090 · Liger-Kernel benchmark scripts held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
TransformersHugging Face Transformers
python
674.8µs
1.02×
Reported
Apache-2.0 · source
LigerLiger-Kernel
triton
659.5µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
xfloat [batch, seq, hidden]
w_gatefloat [intermediate, hidden]
w_upfloat [intermediate, hidden]
w_downfloat [hidden, intermediate]
yfloat [batch, seq, hidden]
Axes and behavior
seqvariable
batchvariable
hiddenvariable
intermediatevariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha2569f5873818070…
Sources: Liger-Kernel benchmarks (2025-11-11) · BSD-2-Clauselast observed 2025-11-11How records are decidedJSON