Skip to content
KernelIndex
Search⌘K

AMD Instinct MI300X

gfx942
Records held
23
Runs
2,099
Operations
4314 implementations

Records held on this GPU

Operation · workload
History
Record
Implementation
Since
2.66ms
2025-09-02
AMD MLA decodebf16 · [128, 1, 7168]
1.78ms
2025-06-03
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 512]
33.1µs
2025-05-25
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 7168] · n 512 · seed 6563
50.1µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 1536]
112.6µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 7168] · n 1536 · seed 6543
214.5µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 2304]
79.4µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 2048]
76.6µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 7168] · n 576 · seed 12346
55.1µs
2025-06-02
Show all 23 records ›
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 7168] · n 1536 · seed 8135
79.0µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 7168] · n 512 · seed 12341
128.2µs
2025-05-25
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 7168] · n 576 · seed 9863
141.1µs
2025-05-25
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 256]
33.7µs
2025-05-27
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 1536]
45.5µs
2025-05-27
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 512]
74.5µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 2304]
308.7µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 7168] · n 4608 · seed 65436
566.3µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 2048]
299.2µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [6144, 256]
77.4µs
2025-06-02
AMD FP8 blockwise GEMMfp32/fp8_e4m3fnuz · [1024, 7168] · n 4608 · seed 7531
142.2µs
2025-06-02

Open in the records ledger →

Coverage by operation family

Family
Operations
Runs
With source
gemm
1
1,764
1,764
moe
1
237
237
mla-attention
1
79
79
trimul
1
19
19

Serving evidence

End-to-end serving results measured on this hardware — a separate corpus with its own resolver, never ranked against the kernel records above.

Model · workload
Scenario
Best reported
Runs
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
24,861 tok/s
4
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
27,854 tok/s
4
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
27,854 tok/s
4
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
24,861 tok/s
4
Llama-2-70B (MLPerf reference)llama2-70b-99 · Interactive
interactive
8,840 tok/s
2
Mixtral-8x7B (MLPerf reference)mixtral-8x7b · Server
server
47,804 tok/s
2
Mixtral-8x7B (MLPerf reference)mixtral-8x7b · Offline
offline
53,334 tok/s
2
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
25,463 tok/s
2

48 serving runs across 28 comparison groups · Open in the serving resolver →

Record activity

No record transitions on this GPU in the trailing year.

Sources: GPU MODE KernelBot (2025-09-23) · June 9 Researcher Reciprocity License v1.0last observed 2025-09-23