Skip to content
KernelIndex
Search⌘K

Trtllm FP8 per tensor scale MoE topk1 e128 h5120 i8192

0 eligible runs
moe

FP8 per-tensor scale MoE (TRT-LLM style). Llama 4 Maverick 17B-128E (EP=1). Llama4 routing. Routing: Llama4 (sigmoid+top-k, routing_method_type=3).

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability

Semantics

Inputs and outputs
routing_logitsbf16 [seq_len, num_experts]
hidden_statesfp8_e4m3 [seq_len, hidden_size]
gemm1_weightsfp8_e4m3 [num_experts, gemm1_out_size, hidden_size]
output1_scales_scalarfp32 [num_experts]
output1_scales_gate_scalarfp32 [num_experts]
gemm2_weightsfp8_e4m3 [num_experts, hidden_size, intermediate_size]
output2_scales_scalarfp32 [num_experts]
routing_biasbf16 [num_experts]
outputbf16 [seq_len, hidden_size]
Axes and behavior
top_kconstant = 1
seq_lenvariable
hidden_sizeconstant = 5120
num_expertsconstant = 128
gemm1_out_sizeconstant = 16384
intermediate_sizeconstant = 8192
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastrtllm_fp8_per_tensor_scale_moe_topk1_e128_h5120_i8192modelllama-4-mavericksha256233703d8b640…
No source imports for this operation yet.How records are decidedJSON