qwen3-next-80b-a3b
best known per operation on the selected GPU
4 operations · 2 families · 20 eligible runs
0 of 4 operations have a usable best known on NVIDIA B200 · 4 without one
Best known on NVIDIA B200
Operation · workload
Best known
Record
Trust
Observed
gqa-ragged-attention1 operation
GQA ragged prefill causal h8 kv1 d256bf16 · [29, 1, 256]flashinfer / wrapper33cb6fnot usable10.2µsReported · Apache-2.0 · source2026-04-09
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 1 more entry in this cohort.
source mirroredApache-2.0no install recipeCohort on the operation page →View source →Vendor the source →Run detail →
Coverage gaps on NVIDIA B200
Operation
Family
Status
gqa-ragged-attention
measured · no usable entry (no install recipe)
No eligible evidence yet for these operations. Challenges →
Related tags: qwen3-next · qwen3-next-80b-a3b-instruct · qwen3-next-80b-a3b-thinking-nvfp4
Sources: FlashInfer-Bench (2026-04-09) · Apache-2.0How comparison worksJSON