Skip to content
KernelIndex
Search⌘K

llama-3-2-3b

best known per operation on the selected GPU
12 operations · 4 families · 79 eligible runs

0 of 12 operations have a usable best known on NVIDIA GB200 · 12 without one

Best known on NVIDIA GB200

Operation · workload
Best known
Record
Trust
Observed
gqa-paged-attention2 operations
87.8µs
Reported · Apache-2.0 · source
2026-04-12

No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1.

39.0µs
Reported · Apache-2.0 · source
2026-04-12

No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1.

gqa-ragged-attention1 operation
41.4µs
Reported · Apache-2.0 · source
2026-04-12

No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1.

Coverage gaps on NVIDIA GB200

Operation
Family
Status
gemm
0 eligible runs · gap
gemm
0 eligible runs · gap
gemm
0 eligible runs · gap
gemm
0 eligible runs · gap
gqa-paged-attention
0 eligible runs on NVIDIA GB200 · measured on NVIDIA B200
gqa-paged-attention
0 eligible runs · gap
sampling
0 eligible runs · gap
sampling
0 eligible runs · gap
sampling
0 eligible runs · gap
gqa-paged-attention
measured · no usable entry (no install recipe)
gqa-paged-attention
measured · no usable entry (no install recipe)
gqa-ragged-attention
measured · no usable entry (no install recipe)

No eligible evidence yet for these operations. Challenges →

Sources: FlashInfer-Bench (2026-04-12) · Apache-2.0How comparison worksJSON