Skip to content
KernelIndex
Search⌘K

GeGLU MLP

84 eligible runs
geglu

GeGLU MLP (tanh-approximated GELU gate) over batch×seq×hidden activations.

Fastest reported · usable today
661.4µsmedian · 1.02× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-11-11. Reported by source; not independently reproduced.

pip install "liger-kernel==0.6.3"

Fastest by languagetriton · Liger · 661.4 µs · #1python · Transformers · 674.3 µs · 1.02×

Current records

Source-native comparison · GPU NVIDIA GeForce RTX 4090 · Workload seq = 1024 · batch = 2 · hidden = 2048 · intermediate = 4096 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 4 results · last observed 2025-11-11Record history →
Estimated floor 58.3 µs · record 11.35× above itestimate, not evidence ›
DRAM 58.3 µs · bandwidth-bound on RTX 4090
every declared tensor crosses HBM exactly once (1,008 GB/s, RTX 4090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
661.4µs
1.00×
Reported · BSD-2-Clause · source
2025-11-11stale

1.02× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
TransformersbaselineHugging Face Transformers
674.3µs
1.02×
Reported · Apache-2.0 · source
2025-11-11stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
3
LigerLiger-Kernel
740.4µs
1.12×
Reported · BSD-2-Clause · source
2025-11-11stale

1.10× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
4
TransformersbaselineHugging Face Transformers
743.4µs
1.12×
Reported · Apache-2.0 · source
2025-11-11stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by seq

5.00 ms10.0 ms1024204840968k16kseq →661.4 µs1.35 ms2.75 ms5.43 ms10.7 msLiger674.3 µs1.41 ms2.82 ms5.70 ms11.3 msTransformers
LigerTransformersbest per workload · NVIDIA GeForce RTX 4090 · Liger-Kernel benchmark scripts held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
TransformersHugging Face Transformers
python
674.3µs
1.02×
Reported
Apache-2.0 · source
LigerLiger-Kernel
triton
661.4µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
xfloat [batch, seq, hidden]
w_gatefloat [intermediate, hidden]
w_upfloat [intermediate, hidden]
w_downfloat [hidden, intermediate]
yfloat [batch, seq, hidden]
Axes and behavior
seqvariable
batchvariable
hiddenvariable
intermediatevariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha25637d2efb3485a…
Sources: Liger-Kernel benchmarks (2025-11-11) · BSD-2-Clauselast observed 2025-11-11How records are decidedJSON