Trtllm fp4 block scale MoE topk10 e128 h2048 i512
0 eligible runs
moe
FP4 block scale MoE (MxFP4+BF16, TRT-LLM style). Qwen3 Next 80B A3B (EP=4, 512/4=128 local experts). Renormalize routing. Routing: Renormalize (TopK->Softmax, routing_method_type=1).
Current records
No published measurement for the selected workload.
Implementations
Implementation
Runtime
Best latency
Evidence
Availability
Semantics
Inputs and outputs
routing_logitsbf16 [seq_len, num_experts]
hidden_statesbf16 [seq_len, hidden_size]
gemm1_weightsfp32 [num_experts, gemm1_out_size, hidden_size]
gemm2_weightsfp32 [num_experts, hidden_size, intermediate_size]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 10
seq_lenvariable
hidden_sizeconstant = 2048
num_expertsconstant = 128
gemm1_out_sizeconstant = 1024
intermediate_sizeconstant = 512
determinismunspecified
constraintsNo mutation or aliasing
Identity
No source imports for this operation yet.How records are decidedJSON