qwen3-14b-nvfp4
0 of 3 operations have a usable best known on NVIDIA B200 · 3 without one
Best known on NVIDIA B200
NVFP4 fused decoder layer pre norm attentionsuite of 16 casesAmir M. Mir | SF Tensornot usable589.1µsReported · license unknown · no source2026-08-20
No entry in this cohort passes the usability policy (no public source, no install recipe, license unknown); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 27 more entries in this cohort.
NVFP4 grouped query attention with KV repeatsuite of 16 casesAmir M. Mir | SF Tensornot usable290.0µsReported · license unknown · no source2026-08-20
No entry in this cohort passes the usability policy (no public source, no install recipe, license unknown); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 27 more entries in this cohort.
Decoder layer full blocksuite of 16 casesac4knot usable877.8µsReported · license unknown · no source2026-07-17
No entry in this cohort passes the usability policy (no public source, no install recipe, license unknown); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 20 more entries in this cohort.
Coverage gaps on NVIDIA B200
No eligible evidence yet for these operations. Challenges →
Related tags: qwen3-14b