Fused neighborhood attention
336 eligible runs
neighborhood-attention
Neighborhood attention module over batch×seq×hidden states: q/k/v projections, banded (kernel_size, dilation) softmax attention, and output projection. The projections are inside the timed module.
Fastest reported · usable today
155.2µsmedian · 67.12× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton
Reported evidence · last observed 2025-05-27. Reported by source; not independently reproduced.
pip install "liger-kernel==0.5.10"Fastest by languagetriton · Liger · 155.2 µs · #1python · PyTorch · 10.4 ms · 67.12×
Current records
Workloadseq = 128 · batch = 2 · heads = 8 · hidden = 512 · dilation = 2 · kernel_size = 7 · fp3228 cases
seq
heads
hidden
batch
dilation
kernel_size
dtype
scrolls · 28 cases total
HardwareBest knownImplementationRuns
NVIDIA GeForce RTX 3090 · env 1155.2 µsLiger2NVIDIA H100 · env 1258.5 µsLiger2NVIDIA H100 · env 2412.2 µsPyTorch2NVIDIA GeForce RTX 3090 · env 2546.1 µsLiger2NVIDIA H100 · env 3890.8 µsLiger2NVIDIA GeForce RTX 3090 · env 32.68 msLiger2Not measured on B200 for this workload. Challenges →
Source-native comparison · GPU NVIDIA GeForce RTX 3090 · Workload seq = 128 · batch = 2 · heads = 8 · hidden = 512 · dilation = 2 · kernel_size = 7 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-05-27Record history →
Estimated floor 5.05 µs · record 30.72× above itestimate, not evidence ›
DRAM 5.05 µs · bandwidth-bound on RTX 3090
every declared tensor crosses HBM exactly once (936 GB/s, RTX 3090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1LigerLiger-Kernel155.2µs1.00×Reported · BSD-2-Clause · source2025-05-27stale
67.12× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
source mirroredBSD-2-ClauseinstallableView source →Run detail →
2PyTorchbaseline10.4ms67.12×Reported · BSD-3-Clause · source2025-05-27stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
source mirroredBSD-3-Clauseno install recipeView source →Run detail →
Scaling by seq
Implementations
Semantics
Inputs and outputs
hidden_statesfloat [batch, seq, hidden]
q_weightfloat [hidden, hidden]
k_weightfloat [hidden, hidden]
v_weightfloat [hidden, hidden]
out_weightfloat [hidden, hidden]
q_biasfloat [hidden]
k_biasfloat [hidden]
v_biasfloat [hidden]
out_biasfloat [hidden]
yfloat [batch, seq, hidden]
Axes and behavior
seqvariable
batchvariable
hiddenvariable
head_dimhead_dim = hidden // heads
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256c262aae772ac…
Sources: Liger-Kernel benchmarks (2025-05-27) · BSD-2-Clauselast observed 2025-05-27How records are decidedJSON