Skip to content
KernelIndex
Search⌘K

Llama 4 rotary position embedding

48 eligible runs
rope

Llama-4-style rotary embedding applied to query and key tensors in batch×seq×heads×head_dim layout (batch 1, complex-frequency formulation).

Fastest reported · usable today
81.3µsmedian · 2.54× faster than baseline
LigerLiger-Kernel · BSD-2-Clause · triton

Reported evidence · last observed 2025-08-07. Reported by source; not independently reproduced.

pip install "liger-kernel==0.6.1"

Fastest by languagetriton · Liger · 81.3 µs · #1python · Transformers · 206.5 µs · 2.54×

Current records

Source-native comparison · GPU NVIDIA H100 · Workload seq = 2048 · hidden = 8192 · q_heads = 32 · kv_heads = 8 · bf16 · Protocol Liger-Kernel benchmark scripts · median · 4 results · last observed 2025-08-07Record history →
Estimated floor 12.5 µs · record 6.49× above itestimate, not evidence ›
DRAM 12.5 µs · bandwidth-bound on H100 SXM
every declared tensor crosses HBM exactly once (3,350 GB/s, H100 SXM datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
LigerLiger-Kernel
81.3µs
1.00×
Reported · BSD-2-Clause · source
2025-08-07stale

2.54× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
2
LigerLiger-Kernel
81.8µs
1.01×
Reported · BSD-2-Clause · source
2025-08-07stale

2.52× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.

source mirroredBSD-2-ClauseinstallableView source →Run detail →
3
TransformersbaselineHugging Face Transformers
206.5µs
2.54×
Reported · Apache-2.0 · source
2025-08-07stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
4
TransformersbaselineHugging Face Transformers
206.6µs
2.54×
Reported · Apache-2.0 · source
2025-08-07stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by seq

500.0 µs1.00 ms1.50 ms1024204840968k16kseq →74.2 µs81.3 µs117.1 µs216.5 µs417.6 µsLiger116.4 µs206.5 µs385.5 µs741.2 µs1.46 msTransformers
LigerTransformersbest per workload · NVIDIA H100 · Liger-Kernel benchmark scripts held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
TransformersHugging Face Transformers
python
37.6µs
1.00×
Reported
Apache-2.0 · source
LigerLiger-Kernel
triton
74.2µs
1.97×
Reported
BSD-2-Clause · source

Semantics

Inputs and outputs
qfloat [1, seq, q_heads, head_dim]
kfloat [1, seq, kv_heads, head_dim]
q_outfloat [1, seq, q_heads, head_dim]
k_outfloat [1, seq, kv_heads, head_dim]
Axes and behavior
seqvariable
q_headsvariable
head_dimhead_dim = hidden // q_heads
kv_headsvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
sha256224b946feeec…
Sources: Liger-Kernel benchmarks (2025-08-07) · BSD-2-Clauselast observed 2025-08-07How records are decidedJSON