llama-4-scout-17b-16e
best known per operation on the selected GPU
3 operations · 2 families · 34 eligible runs
0 of 3 operations have a usable best known on NVIDIA GB200 · 3 without one
Best known on NVIDIA GB200
Operation · workload
Best known
Record
Trust
Observed
gqa-paged-attention2 operations
GQA paged decode h5 kv1 d128 ps1bf16 · [1, 5, 128] · num pages 96 · num kv indices 91flashinfer / wrapper8d9ac8not usable83.4µsReported · Apache-2.0 · source2026-04-12
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1.
source mirroredApache-2.0no install recipeCohort on the operation page →View source →Vendor the source →Run detail →
GQA paged prefill causal h5 kv1 d128 ps1bf16 · [1, 5, 128] · num pages 167 · num kv indices 33flashinfer / wrapper484d2cnot usable105.3µsReproduction-ready · Apache-2.0 · source2026-04-12
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 1 more entry in this cohort.
source mirroredApache-2.0no install recipeCohort on the operation page →View source →Vendor the source →Run detail →
Coverage gaps on NVIDIA GB200
Operation
Family
Status
gqa-paged-attention
measured · no usable entry (no install recipe)
No eligible evidence yet for these operations. Challenges →
Related tags: llama-4-scout
Sources: FlashInfer-Bench (2026-04-12) · Apache-2.0How comparison worksJSON