Multi-token attention
72 eligible runs
multi-token-attention
Multi-token attention over batch×heads_in×seq×seq scores: causal mask, softmax, then a K×K convolution to heads_out with bias, re-masked causally.
Fastest reported · usable today
17.4µsmedian · 14.18× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2025-04-28. Reported by source; not independently reproduced.
pip install "liger-kernel==0.1.1"Fastest by languagetriton · Liger · 17.4 µs · #1python · PyTorch · 246.8 µs · 14.18×
Current records
HardwareBest knownImplementationRuns
NVIDIA GeForce RTX 3090 · env 117.4 µsLiger2NVIDIA GeForce RTX 3090 · env 298.3 µsPyTorch2NVIDIA GeForce RTX 3090 · env 3795.2 µsPyTorch2Not measured on H100, B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA GeForce RTX 3090 · Workload k = 3 · seq = 32 · batch = 2 · heads_in = 4 · heads_out = 4 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-04-30Record history →
Estimated floor 18 ns · record 976.85× above itestimate, not evidence ›
DRAM 18 ns · bandwidth-bound on RTX 3090
every declared tensor crosses HBM exactly once (936 GB/s, RTX 3090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel17.4µs1.00×Reported · BSD-2-Clause · source2025-04-28stale
14.18× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2PyTorchbaseline246.8µs14.18×Reported · BSD-3-Clause · source2025-04-28stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by seq
Implementations
Semantics
Inputs and outputs
scoresfloat [batch, heads_in, seq, seq]
weightfloat [heads_out, heads_in, k, k]
biasfloat [heads_out]
yfloat [batch, heads_out, seq, seq]
Axes and behavior
kvariable
seqvariable
batchvariable
heads_invariable
heads_outvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256daed040c2d77…
Sources: Liger-Kernel benchmarks (2025-04-30) · BSD-2-Clauselast observed 2025-04-30How records are decidedJSON