Skip to content
KernelIndex
Search⌘K

Trtllm FP8 per tensor scale MoE topk8 e128 h2048 i768

0 eligible runs
moe

FP8 per-tensor scale MoE (TRT-LLM style). Qwen3-30B-A3B (EP=1). Renormalize routing (TopK->Softmax). Routing: Renormalize (TopK->Softmax, routing_method_type=1).

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
routing_logitsbf16 [seq_len, num_experts]
hidden_statesfp8_e4m3 [seq_len, hidden_size]
gemm1_weightsfp8_e4m3 [num_experts, gemm1_out_size, hidden_size]
output1_scales_scalarfp32 [num_experts]
output1_scales_gate_scalarfp32 [num_experts]
gemm2_weightsfp8_e4m3 [num_experts, hidden_size, intermediate_size]
output2_scales_scalarfp32 [num_experts]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 8
seq_lenvariable
hidden_sizeconstant = 2048
num_expertsconstant = 128
gemm1_out_sizeconstant = 1536
intermediate_sizeconstant = 768
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastrtllm_fp8_per_tensor_scale_moe_topk8_e128_h2048_i768modelqwen3-30b-a3bsha2565e73fe5a5553…
No source imports for this operation yet.How records are decidedJSON