Skip to content
KernelIndex
Search⌘K

Rotary position embedding

48 eligible runs
rope

Rotary position embedding applied to query and key tensors (batch 1, Llama-style head layout).

Fastest reported · usable today
58.9µsmedian · 8.77× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2024-09-03. Reported by source; not independently reproduced.

pip install "liger-kernel==0.2.1"

Fastest by languagetriton · Liger · 58.9 µs · #1python · Transformers · 516.1 µs · 8.77×

Current records

Source-native comparison · GPU NVIDIA A100 · Workload seq = 2048 · hidden = 8192 · q_heads = 32 · kv_heads = 8 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 4 results · last observed 2024-09-03Record history →
Estimated floor 20.6 µs · record 2.86× above itestimate, not evidence ›
DRAM 20.6 µs · bandwidth-bound on A100 80GB
every declared tensor crosses HBM exactly once (2,039 GB/s, A100 80GB datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
58.9µs
1.00×
Reported · BSD-2-Clause · source
2024-09-03stale

8.77× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
LigerLiger-Kernel
59.5µs
1.01×
Reported · BSD-2-Clause · source
2024-09-03stale

8.68× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
3
TransformersbaselineHugging Face Transformers
516.1µs
8.77×
Reported · Apache-2.0 · source
2024-09-03stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
4
TransformersbaselineHugging Face Transformers
516.8µs
8.78×
Reported · Apache-2.0 · source
2024-09-03stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by seq

100.0 µs1.00 ms1024204840968k16kseq →34.4 µs58.9 µs109.0 µs209.3 µs410.5 µsLiger280.8 µs516.1 µs994.8 µs1.93 ms3.82 msTransformers
LigerTransformersbest per workload · NVIDIA A100 · Liger-Kernel benchmark scripts held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
TransformersHugging Face Transformers
python
79.7µs
7.01×
Reported
Apache-2.0 · source
LigerLiger-Kernel
triton
11.4µs
1.00×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
qfloat [1, q_heads, seq, head_dim]
kfloat [1, kv_heads, seq, head_dim]
q_outfloat [1, q_heads, seq, head_dim]
k_outfloat [1, kv_heads, seq, head_dim]
Axes and behavior
seqvariable
q_headsvariable
head_dimhead_dim = hidden // q_heads
kv_headsvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256e655a9562986…
Sources: Liger-Kernel benchmarks (2024-09-03) · BSD-2-Clauselast observed 2024-09-03How records are decidedJSON