Gdn MTP qk16 v32 d128 k last
0 eligible runs
gdn
Gated Delta Net Multi-Token Prediction (MTP) with GVA configuration. Used for speculative decoding verification where multiple tokens (T > 1) need to be processed in sequence. State layout is k-last [pool_size, H, V, K]. Captured from Qwen3 Next MTP layers.
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
qbf16 [batch_size, seq_len, num_q_heads, head_size]
kbf16 [batch_size, seq_len, num_k_heads, head_size]
vbf16 [batch_size, seq_len, num_v_heads, head_size]
initial_statefp32 [pool_size, num_v_heads, head_size, head_size]
initial_state_indicesint32 [batch_size]
a_logfp32 [num_v_heads]
abf16 [batch_size, seq_len, num_v_heads]
dt_biasfp32 [num_v_heads]
bbf16 [batch_size, seq_len, num_v_heads]
scalefp32 scalar
intermediate_states_bufferfp32 [pool_size, seq_len, num_v_heads, head_size, head_size]
outputbf16 [batch_size, seq_len, num_v_heads, head_size]
final_statefp32 [pool_size, num_v_heads, head_size, head_size]
Axes and behavior
seq_lenvariable
head_sizeconstant = 128
pool_sizevariable
batch_sizevariable
num_k_headsconstant = 16
num_q_headsconstant = 16
num_v_headsconstant = 32
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON