Skip to content
KernelIndex
Search⌘K

MLA ragged prefill causal h8 qk192 vo128

0 eligible runs
mla-ragged

Batched Multi-head Latent Attention prefill with ragged (variable-length) inputs. Uses the absorbed MLA formulation with combined QK dimension (qk_nope=128 + qk_rope=64 = 192) and value output dimension 128. Causal mask is applied. Captured from Kimi K2 / Kimi K2.5 during total prefill (no prefix cache) with tensor parallel size 8 (64/8=8 query heads). Kimi K2.5 shares this shape via its DeepseekV3ForCausalLM text backbone.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
qbf16 [total_q, num_qo_heads, qk_dim]
kbf16 [total_kv, num_kv_heads, qk_dim]
vbf16 [total_kv, num_kv_heads, vo_dim]
qo_indptrint32 [len_indptr]
kv_indptrint32 [len_indptr]
sm_scalefp32 scalar
outputbf16 [total_q, num_qo_heads, vo_dim]
lsefp32 [total_q, num_qo_heads]
Axes and behavior
qk_dimconstant = 192
vo_dimconstant = 128
total_qvariable
total_kvvariable
len_indptrvariable
num_kv_headsconstant = 8
num_qo_headsconstant = 8
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasmla_ragged_prefill_causal_h8_qk192_vo128modelkimi-k2modelkimi-k2-5sha256e50895eef4b9…
No source imports for this operation yet.How records are decidedJSON