Skip to content
KernelIndex
Search⌘K

Dsa sparse attention h16 ckv512 kpe64 topk2048 ps64

0 eligible runs
dsa-paged

Batched Native Sparse Attention (DSA) with sparse TopK KV cache selection. Captured from DeepSeek-V3.2 with tensor parallel size 8. Uses sparse indexing to select only top-K KV cache entries for attention computation. Page size 64 variant. Works for both prefill and decode stages.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
flashinfer / wrapper5af199FlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
q_nopebf16 [num_tokens, num_qo_heads, head_dim_ckv]
q_pebf16 [num_tokens, num_qo_heads, head_dim_kpe]
ckv_cachebf16 [num_pages, page_size, head_dim_ckv]
kpe_cachebf16 [num_pages, page_size, head_dim_kpe]
sparse_indicesint32 [num_tokens, topk]
sm_scalefp32 scalar
outputbf16 [num_tokens, num_qo_heads, head_dim_ckv]
lsefp32 [num_tokens, num_qo_heads]
Axes and behavior
topkconstant = 2048
num_pagesvariable
page_sizeconstant = 64
num_tokensvariable
head_dim_ckvconstant = 512
head_dim_kpeconstant = 64
num_qo_headsconstant = 16
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasdsa_sparse_attention_h16_ckv512_kpe64_topk2048_ps64modeldeepseek-v3-2sha2568b169399c884…
No source imports for this operation yet.How records are decidedJSON