Skip to content
KernelIndex
Search⌘K

Multi-token attention

72 eligible runs
multi-token-attention

Multi-token attention over batch×heads_in×seq×seq scores: causal mask, softmax, then a K×K convolution to heads_out with bias, re-masked causally.

Fastest reported · usable today
17.4µsmedian · 14.18× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-04-28. Reported by source; not independently reproduced.

pip install "liger-kernel==0.1.1"

Fastest by languagetriton · Liger · 17.4 µs · #1python · PyTorch · 246.8 µs · 14.18×

Current records

Workloadk = 3 · seq = 32 · batch = 2 · heads_in = 4 · heads_out = 4 · bf16k = 3 · batch = 2 · heads_in = 4 · heads_out = 412 cases

Not measured on H100, B200 for this workload. Challenges →

Source-native comparison · GPU NVIDIA GeForce RTX 3090 · Workload k = 3 · seq = 32 · batch = 2 · heads_in = 4 · heads_out = 4 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-04-30Record history →
Estimated floor 18 ns · record 976.85× above itestimate, not evidence ›
DRAM 18 ns · bandwidth-bound on RTX 3090
every declared tensor crosses HBM exactly once (936 GB/s, RTX 3090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
17.4µs
1.00×
Reported · BSD-2-Clause · source
2025-04-28stale

14.18× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
PyTorchbaseline
246.8µs
14.18×
Reported · BSD-3-Clause · source
2025-04-28stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →

Scaling by seq

100.0 µs32641282565121024seq →17.4 µs18.4 µs23.6 µs43.0 µs126.0 µs528.4 µsLiger246.8 µs241.7 µs242.7 µs241.7 µs313.3 µs719.9 µsPyTorch
LigerPyTorchbest per workload · NVIDIA GeForce RTX 3090 · Liger-Kernel benchmark scripts held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
63.5µs
3.65×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
17.4µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
scoresfloat [batch, heads_in, seq, seq]
weightfloat [heads_out, heads_in, k, k]
biasfloat [heads_out]
yfloat [batch, heads_out, seq, seq]
Axes and behavior
kvariable
seqvariable
batchvariable
heads_invariable
heads_outvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256daed040c2d77…
Sources: Liger-Kernel benchmarks (2025-04-30) · BSD-2-Clauselast observed 2025-04-30How records are decidedJSON