Skip to content
KernelIndex
Search⌘K

RoPE with cos sin cache neox style d128 rd64

0 eligible runs
rope

Rotary Position Embedding (RoPE) with pre-computed cos/sin cache, NeoX-style interleaving, and partial rotary dimension. head_size=128, rotary_dim=64. NeoX style splits the rotary dimensions into two halves [x1, x2] and applies rotation, as opposed to GPT-J style which interleaves even/odd indices. Only the first 64 dimensions are rotated; the remaining 64 pass through unchanged. Matches the FlashInfer API flashinfer.rope.apply_rope_with_cos_sin_cache_inplace. Captured from MiniMax M2.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
qbf16 [num_tokens, num_qo_heads, head_size]
kbf16 [num_tokens, num_kv_heads, head_size]
cos_sin_cachefp32 [max_seq_len, rotary_dim]
positionsint64 [num_tokens]
q_outbf16 [num_tokens, num_qo_heads, head_size]
k_outbf16 [num_tokens, num_kv_heads, head_size]
Axes and behavior
head_sizeconstant = 128
num_tokensvariable
rotary_dimconstant = 64
max_seq_lenvariable
num_kv_headsvariable
num_qo_headsvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasrope_with_cos_sin_cache_neox_style_d128_rd64modelminimax-m2sha256b6d340cf8a69…
No source imports for this operation yet.How records are decidedJSON