Dsa sparse attention h16 ckv512 kpe64 topk2048 ps64
0 eligible runs
dsa-paged
Batched Native Sparse Attention (DSA) with sparse TopK KV cache selection. Captured from DeepSeek-V3.2 with tensor parallel size 8. Uses sparse indexing to select only top-K KV cache entries for attention computation. Page size 64 variant. Works for both prefill and decode stages.
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
q_nopebf16 [num_tokens, num_qo_heads, head_dim_ckv]
q_pebf16 [num_tokens, num_qo_heads, head_dim_kpe]
ckv_cachebf16 [num_pages, page_size, head_dim_ckv]
kpe_cachebf16 [num_pages, page_size, head_dim_kpe]
sparse_indicesint32 [num_tokens, topk]
sm_scalefp32 scalar
outputbf16 [num_tokens, num_qo_heads, head_dim_ckv]
lsefp32 [num_tokens, num_qo_heads]
Axes and behavior
topkconstant = 2048
num_pagesvariable
page_sizeconstant = 64
num_tokensvariable
head_dim_ckvconstant = 512
head_dim_kpeconstant = 64
num_qo_headsconstant = 16
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON