Skip to content
KernelIndex
Search⌘K

Fused neighborhood attention

336 eligible runs
neighborhood-attention

Neighborhood attention module over batch×seq×hidden states: q/k/v projections, banded (kernel_size, dilation) softmax attention, and output projection. The projections are inside the timed module.

Fastest reported · usable today
155.2µsmedian · 67.12× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-05-27. Reported by source; not independently reproduced.

pip install "liger-kernel==0.5.10"

Fastest by languagetriton · Liger · 155.2 µs · #1python · PyTorch · 10.4 ms · 67.12×

Current records

Source-native comparison · GPU NVIDIA GeForce RTX 3090 · Workload seq = 128 · batch = 2 · heads = 8 · hidden = 512 · dilation = 2 · kernel_size = 7 · fp32 · Protocol Liger-Kernel benchmark scripts · median · 2 results · last observed 2025-05-27Record history →
Estimated floor 5.05 µs · record 30.72× above itestimate, not evidence ›
DRAM 5.05 µs · bandwidth-bound on RTX 3090
every declared tensor crosses HBM exactly once (936 GB/s, RTX 3090 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
155.2µs
1.00×
Reported · BSD-2-Clause · source
2025-05-27stale

67.12× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
PyTorchbaseline
10.4ms
67.12×
Reported · BSD-3-Clause · source
2025-05-27stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredBSD-3-Clauseno install recipeView source →Run detail →

Scaling by seq

1.00 ms10.0 ms100.0 ms64128256512102420484096seq →170.0 µs155.2 µs170.0 µs330.8 µs855.0 µs2.37 ms8.25 msLiger5.06 ms10.4 ms21.1 ms39.9 ms87.5 ms162.8 ms318.9 msPyTorch
LigerPyTorchbest per workload · NVIDIA GeForce RTX 3090 · Liger-Kernel benchmark scripts held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
python
410.8µs
2.65×
Reported
BSD-3-Clause · source
LigerLiger-Kernel
triton
155.2µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
hidden_statesfloat [batch, seq, hidden]
q_weightfloat [hidden, hidden]
k_weightfloat [hidden, hidden]
v_weightfloat [hidden, hidden]
out_weightfloat [hidden, hidden]
q_biasfloat [hidden]
k_biasfloat [hidden]
v_biasfloat [hidden]
out_biasfloat [hidden]
yfloat [batch, seq, hidden]
Axes and behavior
seqvariable
batchvariable
hiddenvariable
head_dimhead_dim = hidden // heads
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256c262aae772ac…
Sources: Liger-Kernel benchmarks (2025-05-27) · BSD-2-Clauselast observed 2025-05-27How records are decidedJSON