Sparsemax
36 eligible runs
sparsemax
Sparsemax projection over the last dimension of a tokens×features tensor.
Fastest reported · usable today
44.0µsmedian · 4.12× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2025-04-28. Reported by source; not independently reproduced.
pip install "liger-kernel==0.5.8"Fastest by languagetriton · Liger · 44.0 µs · #1python · PyTorch · 181.2 µs · 4.12×
Current records
Workloadtokens = 2048 · features = 4096 · fp32tokens = 2048 · features = 8192 · fp32tokens = 2048 · features = 16384 · fp32tokens = 2048 · features = 32768 · fp32tokens = 2048 · features = 1024 · fp32tokens = 2048 · features = 2048 · fp32
HardwareBest knownImplementationRuns
NVIDIA GeForce RTX 3090 · env 144.0 µsLiger2NVIDIA GeForce RTX 3090 · env 2414.7 µsLiger2NVIDIA GeForce RTX 3090 · env 3455.9 µsLiger2Not measured on H100, B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA GeForce RTX 3090 · Workload tokens = 2048 · features = 1024 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-04-28Record history →
Estimated floor 8.96 µs · record 4.91× above itestimate, not evidence ›
DRAM 8.96 µs · bandwidth-bound on RTX 3090
every declared tensor crosses HBM exactly once (936 GB/s, RTX 3090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel44.0µs1.00×Reported · BSD-2-Clause · source2025-04-28stale
4.12× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2PyTorchbaseline181.2µs4.12×Reported · BSD-3-Clause · source2025-04-28stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by features
Implementations
Semantics
Inputs and outputs
xfloat [tokens, features]
yfloat [tokens, features]
Axes and behavior
tokensvariable
featuresvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256f124d7a3760e…
Sources: Liger-Kernel benchmarks (2025-04-28) · BSD-2-Clauselast observed 2025-04-28How records are decidedJSON