Llama 4 rotary position embedding
Llama-4-style rotary embedding applied to query and key tensors in batch×seq×heads×head_dim layout (batch 1, complex-frequency formulation).
Reported evidence · last observed 2025-08-07. Reported by source; not independently reproduced.
pip install "liger-kernel==0.6.1"Fastest by languagetriton · Liger · 81.3 µs · #1python · Transformers · 206.5 µs · 2.54×
Current records
Not measured on B200 for this workload. Challenges →
Estimated floor 12.5 µs · record 6.49× above itestimate, not evidence ›
1LigerLiger-Kernel81.3µs1.00×Reported · BSD-2-Clause · source2025-08-07stale
2.54× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
2LigerLiger-Kernel81.8µs1.01×Reported · BSD-2-Clause · source2025-08-07stale
2.52× faster than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.
3206.5µs2.54×Reported · Apache-2.0 · source2025-08-07stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
4206.6µs2.54×Reported · Apache-2.0 · source2025-08-07stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.