Skip to content
KernelIndex
Search⌘K

Mamba ssu decode h64 d64 s128 ng4

0 eligible runs
mamba-ssu

Mamba2 Selective State Update (SSU) decode with 64 heads, head_dim=64, dstate=128, 4 groups. Single-token generation via recurrent SSM state update. Captured from NVIDIA NemotronH-8B Mamba layers (TP=2, heads split by TP: 128/2=64, groups split by TP: 8/2=4). nheads/ngroups=16 is supported by FlashInfer.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
statebf16 [state_cache_size, nheads, head_dim, dstate]
xbf16 [batch_size, nheads, head_dim]
dtbf16 [batch_size, nheads, head_dim]
afp32 [nheads, head_dim, dstate]
bbf16 [batch_size, ngroups, dstate]
cbf16 [batch_size, ngroups, dstate]
dbf16 [nheads, head_dim]
dt_biasbf16 [nheads, head_dim]
state_batch_indicesint32 [batch_size]
outputbf16 [batch_size, nheads, head_dim]
new_statebf16 [state_cache_size, nheads, head_dim, dstate]
Axes and behavior
dstateconstant = 128
nheadsconstant = 64
ngroupsconstant = 4
head_dimconstant = 64
batch_sizevariable
state_cache_sizevariable
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasmamba_ssu_decode_h64_d64_s128_ng4modelnemotron-h-8bsha256a7ce82384e28…
No source imports for this operation yet.How records are decidedJSON