Dsa topk indexer FP8 h64 d128 topk2048 ps64
0 eligible runs
dsa-paged
Native Sparse Attention (DSA) TopK indexer with FP8 quantization for DeepSeek-V3.2. Computes sparse attention scores using ReLU activation and learned weights, then selects top-K KV cache indices. Formula: sum(relu(q @ K.T) * weights). Matches SGLang/deep_gemm implementation. Page size 64 variant.
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
q_index_fp8fp8_e4m3 [batch_size, num_index_heads, index_head_dim]
k_index_cache_fp8int8 [num_pages, page_size, kv_cache_num_heads, head_dim_with_scale]
weightsfp32 [batch_size, num_index_heads]
seq_lensint32 [batch_size]
block_tableint32 [batch_size, max_num_pages]
topk_indicesint32 [batch_size, topk]
Axes and behavior
topkconstant = 2048
num_pagesvariable
page_sizeconstant = 64
batch_sizevariable
max_num_pagesvariable
index_head_dimconstant = 128
num_index_headsconstant = 64
kv_cache_num_headsconstant = 1
head_dim_with_scaleconstant = 132
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON