llama-3-2-3b
best known per operation on the selected GPU
12 operations · 4 families · 79 eligible runs
0 of 12 operations have a usable best known on NVIDIA B200 · 12 without one
Best known on NVIDIA B200
Operation · workload
Best known
Record
Trust
Observed
gqa-paged-attention1 operation
GQA paged decode h24 kv8 d128 ps1bf16 · [1, 24, 128] · num pages 14 · num kv indices 13flashinfer / wrapper96864enot usable20.4µsReported · Apache-2.0 · source2026-04-09
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1.
source mirroredApache-2.0no install recipeCohort on the operation page →View source →Vendor the source →Run detail →
Coverage gaps on NVIDIA B200
Operation
Family
Status
gqa-paged-attention
0 eligible runs on NVIDIA B200 · measured on NVIDIA GB200
gqa-paged-attention
0 eligible runs on NVIDIA B200 · measured on NVIDIA GB200
gqa-ragged-attention
0 eligible runs on NVIDIA B200 · measured on NVIDIA GB200
No eligible evidence yet for these operations. Challenges →
Sources: FlashInfer-Bench (2026-04-12) · Apache-2.0How comparison worksJSON