Skip to content
KernelIndex
Search⌘K

Trtllm fp4 block scale routed MoE topk10 e128 h2048 i512

0 eligible runs
moe

FP4 block scale routed MoE (MxFP4+BF16, TRT-LLM style). Qwen3 Next 80B A3B (EP=4, 512/4=128 local experts). Renormalize routing. Pre-computed routing. Routing: Renormalize (pre-computed topk_ids, routing_method_type=1).

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
topk_idsint32 [seq_len, top_k]
hidden_statesbf16 [seq_len, hidden_size]
gemm1_weightsfp32 [num_experts, gemm1_out_size, hidden_size]
gemm2_weightsfp32 [num_experts, hidden_size, intermediate_size]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 10
seq_lenvariable
hidden_sizeconstant = 2048
num_expertsconstant = 128
gemm1_out_sizeconstant = 1024
intermediate_sizeconstant = 512
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastrtllm_fp4_block_scale_routed_moe_topk10_e128_h2048_i512modelqwen3-next-80b-a3bsha2564d8ca7919539…
No source imports for this operation yet.How records are decidedJSON