Gdn prefill qk4 v8 d128 k last
0 eligible runs
gdn
Gated Delta Net prefill with GVA configuration and k-last state layout. The state is in k-last layout [N, H, V, K]. Captured from Qwen3 Next linear attention layers (TP=4).
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
qbf16 [total_seq_len, num_q_heads, head_size]
kbf16 [total_seq_len, num_k_heads, head_size]
vbf16 [total_seq_len, num_v_heads, head_size]
statefp32 [num_seqs, num_v_heads, head_size, head_size]
a_logfp32 [num_v_heads]
abf16 [total_seq_len, num_v_heads]
dt_biasfp32 [num_v_heads]
bbf16 [total_seq_len, num_v_heads]
cu_seqlensint64 [len_cu_seqlens]
scalefp32 scalar
outputbf16 [total_seq_len, num_v_heads, head_size]
new_statefp32 [num_seqs, num_v_heads, head_size, head_size]
Axes and behavior
num_seqsvariable
head_sizeconstant = 128
num_k_headsconstant = 4
num_q_headsconstant = 4
num_v_headsconstant = 8
total_seq_lenvariable
len_cu_seqlensvariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON