Skip to content
KernelIndex
Search⌘K

Mamba chunk scan with segsum

20 eligible runs
ssm

Mamba-2 chunk-based parallel scan with segment sum computation for intra-chunk and inter-chunk state propagation. This is the core parallel scanning algorithm that enables efficient SSM computation by breaking sequences into chunks (chunk_size=256), computing intra-chunk attention-like patterns using segment sums, maintaining inter-chunk state recurrence with exponential decay, and combining diagonal (intra-chunk) and off-diagonal (inter-chunk) contributions.

Fastest reported
48.5µsmean · 93% of SOL
Amir M. Mir | SF TensorLicense unknown · unknown

Reported evidence · last observed 2026-07-10. Reported by source; not independently reproduced. No public source. License unknown.

No public sourceRun detail →

Current records

Workloadsuite of 16 cases · mean latencyhead_dim = 64 · n_groups = 1 · num_heads = 16 · chunk_size = 256 · hidden_out = 1024 · state_size = 25616 cases

Not measured on H100 for this workload. Challenges →

Source-native comparison · GPU NVIDIA B200 · Workload suite of 16 cases · mean latency · Protocol SOL-ExecBench evaluation stack v1.1 · mean · 12 results · last observed 2026-08-15Record history →
#
Implementation
Latency
vs #1
Trust
Observed
1
48.5µs93% SOL
1.00×
Reported · license unknown · no source
2026-07-10

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
2
60.8µs92% SOL
1.25×
Reported · license unknown · no source
2026-07-10

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
3
62.5µs92% SOL
1.29×
Reported · license unknown · no source
2026-07-10

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
4
64.9µs91% SOL
1.34×
Reported · license unknown · no source
2026-08-13

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
5
68.0µs91% SOL
1.40×
Reported · license unknown · no source
2026-07-10

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
6
73.9µs90% SOL
1.52×
Reported · license unknown · no source
2026-07-21

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
7
84.6µs89% SOL
1.74×
Reported · license unknown · no source
2026-07-17

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
8
99.0µs87% SOL
2.04×
Reported · license unknown · no source
2026-08-15

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
9
99.1µs87% SOL
2.04×
Reported · license unknown · no source
2026-08-15

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
10
99.1µs87% SOL
2.04×
Reported · license unknown · no source
2026-08-15

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
11
107.2µs86% SOL
2.21×
Reported · license unknown · no source
2026-07-26

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →
12
473.4µs58% SOL
9.75×
Reported · license unknown · no source
2026-07-10

Measured exactly what you asked. Reported by source; not independently reproduced.

no sourcelicense unknownno install recipeRun detail →

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
unknown
58.9µs
1.25×
Reported
license unknown · no source
unknown
58.8µs
1.25×
Reported
license unknown · no source
unknown
64.9µs
1.38×
Reported
license unknown · no source
unknown
73.9µs
1.58×
Reported
license unknown · no source
unknown
84.6µs
1.80×
Reported
license unknown · no source
unknown
99.0µs
2.11×
Reported
license unknown · no source
unknown
99.1µs
2.11×
Reported
license unknown · no source
unknown
99.1µs
2.11×
Reported
license unknown · no source
Show all 20 implementations ›
unknown
473.4µs
10.09×
Reported
license unknown · no source
unknown
60.8µs
1.30×
Reported
license unknown · no source
unknown
62.5µs
1.33×
Reported
license unknown · no source
unknown
68.0µs
1.45×
Reported
license unknown · no source
unknown
107.2µs
2.28×
Reported
license unknown · no source
unknown
152.4µs
3.25×
Reported
license unknown · no source
unknown
493.6µs
10.52×
Reported
license unknown · no source
unknown
565.2µs
12.04×
Reported
license unknown · no source
unknown
584.3µs
12.45×
Reported
license unknown · no source
unknown
46.9µs
1.00×
Reported
license unknown · no source
unknown
48.5µs
1.03×
Reported
license unknown · no source
unknown
57.6µs
1.23×
Reported
license unknown · no source

Semantics

Inputs and outputs
hidden_statesbf16 [batch_size, seq_len, num_heads, head_dim]
abf16 [batch_size, num_heads, seq_len]
bbf16 [batch_size, seq_len, n_groups, state_size]
cbf16 [batch_size, seq_len, n_groups, state_size]
dbf16 [num_heads]
initial_statesbf16 [batch_size, num_heads, head_dim, state_size]
outputbf16 [batch_size, seq_len, hidden_out]
final_statebf16 [batch_size, num_heads, head_dim, state_size]
Axes and behavior
seq_lenvariable
head_dimconstant = 64
n_groupsconstant = 1
num_headsconstant = 16
batch_sizevariable
chunk_sizeconstant = 256
hidden_outhidden_out = num_heads * head_dim
state_sizeconstant = 256
determinismunspecified
constraintsNo mutation or aliasing
Identity
alias043_mamba_chunk_scan_with_segsummodelgranite-4-0-1bsha256df0403adcba7…
Sources: NVIDIA SOL-ExecBench (2026-08-15)last observed 2026-08-15How records are decidedJSON