Skip to content
KernelIndex
Search⌘K

Trtllm fp4 block scale MoE topk8 e64 h4096 i1536

0 eligible runs
moe

FP4 block scale MoE (MxFP4+BF16, TRT-LLM style). Qwen3-235B-A22B (EP=2, 128/2=64 local experts). Renormalize routing. Routing: Renormalize (TopK->Softmax, routing_method_type=1).

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
routing_logitsbf16 [seq_len, num_experts]
hidden_statesbf16 [seq_len, hidden_size]
gemm1_weightsfp32 [num_experts, gemm1_out_size, hidden_size]
gemm2_weightsfp32 [num_experts, hidden_size, intermediate_size]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 8
seq_lenvariable
hidden_sizeconstant = 4096
num_expertsconstant = 64
gemm1_out_sizeconstant = 3072
intermediate_sizeconstant = 1536
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastrtllm_fp4_block_scale_moe_topk8_e64_h4096_i1536modelqwen3-235b-a22bsha256056e39950cef…
No source imports for this operation yet.How records are decidedJSON