MoE FP8 block scale renorm topk10 e128 h2048 i512
0 eligible runs
moe
FP8 block-scale MoE (DeepSeek-style). Qwen3-Next 80B A3B (EP=1). Renormalize routing (TopK->Softmax).
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
routing_logitsbf16 [seq_len, num_local_experts]
hidden_statesfp8_e4m3 [seq_len, hidden_size]
hidden_states_scalefp32 [num_hidden_blocks, seq_len]
gemm1_weightsfp8_e4m3 [num_local_experts, gemm1_out_size, hidden_size]
gemm1_weights_scalefp32 [num_local_experts, num_gemm1_out_blocks, num_hidden_blocks]
gemm2_weightsfp8_e4m3 [num_local_experts, hidden_size, intermediate_size]
gemm2_weights_scalefp32 [num_local_experts, num_hidden_blocks, num_intermediate_blocks]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 10
seq_lenvariable
hidden_sizeconstant = 2048
gemm1_out_sizeconstant = 1024
intermediate_sizeconstant = 512
num_hidden_blocksconstant = 16
num_local_expertsconstant = 128
num_gemm1_out_blocksconstant = 8
num_intermediate_blocksconstant = 4
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON