MoE FP8 block scale renorm topk8 e128 h4096 i1536
0 eligible runs
moe
FP8 block-scale MoE (DeepSeek-style). Qwen3-235B-A22B (EP=1). Renormalize routing (TopK->Softmax).
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
routing_logitsbf16 [seq_len, num_local_experts]
hidden_statesfp8_e4m3 [seq_len, hidden_size]
hidden_states_scalefp32 [num_hidden_blocks, seq_len]
gemm1_weightsfp8_e4m3 [num_local_experts, gemm1_out_size, hidden_size]
gemm1_weights_scalefp32 [num_local_experts, num_gemm1_out_blocks, num_hidden_blocks]
gemm2_weightsfp8_e4m3 [num_local_experts, hidden_size, intermediate_size]
gemm2_weights_scalefp32 [num_local_experts, num_hidden_blocks, num_intermediate_blocks]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 8
seq_lenvariable
hidden_sizeconstant = 4096
gemm1_out_sizeconstant = 3072
intermediate_sizeconstant = 1536
num_hidden_blocksconstant = 32
num_local_expertsconstant = 128
num_gemm1_out_blocksconstant = 24
num_intermediate_blocksconstant = 12
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON