llama-3-1-405b
best known per operation on the selected GPU
4 operations · 2 families · 39 eligible runs
0 of 4 operations have a usable best known on NVIDIA B200 · 4 without one
Best known on NVIDIA B200
Operation · workload
Best known
Record
Trust
Observed
attention1 operation
Fused attention QK matmul scale mask softmax backwardsuite of 16 casesAmir M. Mir | SF Tensornot usable137.0µsReported · license unknown · no source2026-08-22
No entry in this cohort passes the usability policy (no public source, no install recipe, license unknown); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 26 more entries in this cohort.
no sourcelicense unknownno install recipeCohort on the operation page →Run detail →
Coverage gaps on NVIDIA B200
Operation
Family
Status
attention
measured · no usable entry (no public source, no install recipe, license unknown)
No eligible evidence yet for these operations. Challenges →
Sources: NVIDIA SOL-ExecBench (2026-08-24)How comparison worksJSON