MLA ragged prefill causal h16 qk192 vo128
Batched Multi-head Latent Attention prefill with ragged (variable-length) inputs. Uses the absorbed MLA formulation with combined QK dimension (qk_nope=128 + qk_rope=64 = 192) and value output dimension 128. Causal mask is applied. Captured from DeepSeek-V3 during total prefill (no prefix cache) with tensor parallel size 8.
Reported evidence · last observed 2026-04-06. The source's designated baseline implementation. Reported by source; not independently reproduced.
Current records
Workloadtotal_q = 86 · total_kv = 86 · len_indptr = 3 · bf1610 cases
Not measured on H100 for this workload. Challenges →
Estimated floor 176 ns · record 58.08× above itestimate, not evidence ›
110.2µs1.00×Reported · Apache-2.0 · source2026-04-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
210.3µs1.01×Reported · Apache-2.0 · source2026-04-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.