Skip to content
KernelIndex
Search⌘K

Softmax

48 eligible runs
softmax

Softmax over the last dimension of a rows×hidden tensor.

Best usable
7.17µsmedian · 1.17× vs fastest
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-04-30. Reported by source; not independently reproduced.

pip install "liger-kernel==0.5.8"
Source baseline · unbeaten · not usable as-is
6.14µsmedian
PyTorchBSD-3-Clause · python

Reported evidence · last observed 2025-04-30. The source's designated baseline implementation. Reported by source; not independently reproduced.

Fastest by languagepython · PyTorch · 6.14 µs · #1triton · Liger · 7.17 µs · 1.17×

Current records

Workloadrows = 2048 · hidden = 128 · fp32rows = 204812 cases

Not measured on H100, B200 for this workload. Challenges →

Source-native comparison · GPU NVIDIA GeForce RTX 3090 · Workload rows = 2048 · hidden = 128 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-04-30Record history →
Estimated floor 1.12 µs · record 5.48× above itestimate, not evidence ›
DRAM 1.12 µs · bandwidth-bound on RTX 3090
every declared tensor crosses HBM exactly once (936 GB/s, RTX 3090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
PyTorchbaseline
6.14µs
1.00×
Reported · BSD-3-Clause · source
2025-04-30stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →
2
LigerLiger-Kernel
7.17µs
1.17×
Reported · BSD-2-Clause · source
2025-04-30stale

1.17× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →

Scaling by hidden

20.0 µs40.0 µs60.0 µs80.0 µs128256512102420484096hidden →7.17 µs8.45 µs13.3 µs21.5 µs41.0 µs79.9 µsLiger6.14 µs8.19 µs12.3 µs22.5 µs57.6 µs83.2 µsPyTorch
LigerPyTorchbest per workload · NVIDIA GeForce RTX 3090 · Liger-Kernel benchmark scripts held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
6.14µs
1.00×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
6.34µs
1.03×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
xfloat [rows, hidden]
yfloat [rows, hidden]
Axes and behavior
rowsvariable
hiddenvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha25648e06011fd0c…
Sources: Liger-Kernel benchmarks (2025-04-30) · BSD-2-Clauselast observed 2025-04-30How records are decidedJSON