Skip to content
KernelIndex
Search⌘K

MoE FP8 block scale renorm topk10 e128 h2048 i512

0 eligible runs
moe

FP8 block-scale MoE (DeepSeek-style). Qwen3-Next 80B A3B (EP=1). Renormalize routing (TopK->Softmax).

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
routing_logitsbf16 [seq_len, num_local_experts]
hidden_statesfp8_e4m3 [seq_len, hidden_size]
hidden_states_scalefp32 [num_hidden_blocks, seq_len]
gemm1_weightsfp8_e4m3 [num_local_experts, gemm1_out_size, hidden_size]
gemm1_weights_scalefp32 [num_local_experts, num_gemm1_out_blocks, num_hidden_blocks]
gemm2_weightsfp8_e4m3 [num_local_experts, hidden_size, intermediate_size]
gemm2_weights_scalefp32 [num_local_experts, num_hidden_blocks, num_intermediate_blocks]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 10
seq_lenvariable
hidden_sizeconstant = 2048
gemm1_out_sizeconstant = 1024
intermediate_sizeconstant = 512
num_hidden_blocksconstant = 16
num_local_expertsconstant = 128
num_gemm1_out_blocksconstant = 8
num_intermediate_blocksconstant = 4
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasmoe_fp8_block_scale_renorm_topk10_e128_h2048_i512modelqwen3-next-80b-a3bsha25640efdb71d428…
No source imports for this operation yet.How records are decidedJSON