MLA ragged prefill causal h8 qk192 vo128
0 eligible runs
mla-ragged
Batched Multi-head Latent Attention prefill with ragged (variable-length) inputs. Uses the absorbed MLA formulation with combined QK dimension (qk_nope=128 + qk_rope=64 = 192) and value output dimension 128. Causal mask is applied. Captured from Kimi K2 / Kimi K2.5 during total prefill (no prefix cache) with tensor parallel size 8 (64/8=8 query heads). Kimi K2.5 shares this shape via its DeepseekV3ForCausalLM text backbone.
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
qbf16 [total_q, num_qo_heads, qk_dim]
kbf16 [total_kv, num_kv_heads, qk_dim]
vbf16 [total_kv, num_kv_heads, vo_dim]
qo_indptrint32 [len_indptr]
kv_indptrint32 [len_indptr]
sm_scalefp32 scalar
outputbf16 [total_q, num_qo_heads, vo_dim]
lsefp32 [total_q, num_qo_heads]
Axes and behavior
qk_dimconstant = 192
vo_dimconstant = 128
total_qvariable
total_kvvariable
len_indptrvariable
num_kv_headsconstant = 8
num_qo_headsconstant = 8
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON