Skip to content
KernelIndex
Search⌘K

MLA paged prefill causal h8 ckv512 kpe64 ps1

0 eligible runs
mla-paged-attention

Batched Multi-head Latent Attention prefill with a paged KV cache. Causal mask is applied. Captured from Kimi K2 / Kimi K2.5 during incremental prefill with tensor parallel size 8 (64/8=8 query heads). Kimi K2.5 shares this shape via its DeepseekV3ForCausalLM text backbone.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
q_nopebf16 [total_q, num_qo_heads, head_dim_ckv]
q_pebf16 [total_q, num_qo_heads, head_dim_kpe]
ckv_cachebf16 [num_pages, page_size, head_dim_ckv]
kpe_cachebf16 [num_pages, page_size, head_dim_kpe]
qo_indptrint32 [len_indptr]
kv_indptrint32 [len_indptr]
kv_indicesint32 [num_kv_indices]
sm_scalefp32 scalar
outputbf16 [total_q, num_qo_heads, head_dim_ckv]
lsefp32 [total_q, num_qo_heads]
Axes and behavior
total_qvariable
num_pagesvariable
page_sizeconstant = 1
len_indptrvariable
head_dim_ckvconstant = 512
head_dim_kpeconstant = 64
num_qo_headsconstant = 8
num_kv_indicesvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasmla_paged_prefill_causal_h8_ckv512_kpe64_ps1modelkimi-k2modelkimi-k2-5sha25651665866d449…
No source imports for this operation yet.How records are decidedJSON