Skip to content
KernelIndex
Search⌘K

Gdn decode qk8 v16 d128 k last

0 eligible runs
gdn

Gated Delta Net decode with GVA configuration and k-last state layout. Single-token generation with recurrent state update. Captured from Qwen3 Next linear attention layers (TP=2).

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
flashinfer / wrappera5e9d2FlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
qbf16 [batch_size, seq_len, num_q_heads, head_size]
kbf16 [batch_size, seq_len, num_k_heads, head_size]
vbf16 [batch_size, seq_len, num_v_heads, head_size]
statefp32 [batch_size, num_v_heads, head_size, head_size]
a_logfp32 [num_v_heads]
abf16 [batch_size, seq_len, num_v_heads]
dt_biasfp32 [num_v_heads]
bbf16 [batch_size, seq_len, num_v_heads]
scalefp32 scalar
outputbf16 [batch_size, seq_len, num_v_heads, head_size]
new_statefp32 [batch_size, num_v_heads, head_size, head_size]
Axes and behavior
seq_lenconstant = 1
head_sizeconstant = 128
batch_sizevariable
num_k_headsconstant = 8
num_q_headsconstant = 8
num_v_headsconstant = 16
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasgdn_decode_qk8_v16_d128_k_lastmodelqwen3-nextmodelqwen3-5-35b-a3bsha256cea2d0f39fe7…
No source imports for this operation yet.How records are decidedJSON