Skip to content
KernelIndex
Search⌘K

torch / matmul075b0d

torch_matmul_075b0d · FlashInfer-Bench baselines · python · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 7 lines, Apache-2.0, pinned at da91508.

main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-torch-matmul-075b0d?include=source"
interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
measured onNVIDIA B200
declared hardwareCPU, NVIDIA A100, NVIDIA B200, NVIDIA H100
architecturessm_100, sm_90, unknown
dtypesfp16

Benchmark evidence

25 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n5120 k2048fp16 · [6, 2048]
NVIDIA B200
12.0µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [8, 2048]
NVIDIA B200
12.1µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [2, 2048]
NVIDIA B200
12.2µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [4, 2048]
NVIDIA B200
12.3µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [5, 2048]
NVIDIA B200
12.3µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [1, 2048]
NVIDIA B200
12.3µs
#1= of 6
2025-10-20
GEMM n5120 k2048fp16 · [32, 2048]
NVIDIA B200
12.3µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [64, 2048]
NVIDIA B200
12.4µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [16, 2048]
NVIDIA B200
12.5µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [63, 2048]
NVIDIA B200
12.6µs
#2 of 6
2025-10-20
Show all 25 measurements ›
GEMM n5120 k2048fp16 · [17, 2048]
NVIDIA B200
12.8µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [25, 2048]
NVIDIA B200
12.9µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [128, 2048]
NVIDIA B200
12.9µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [93, 2048]
NVIDIA B200
13.0µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [34, 2048]
NVIDIA B200
13.0µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [172, 2048]
NVIDIA B200
15.2µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [492, 2048]
NVIDIA B200
15.9µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [289, 2048]
NVIDIA B200
16.2µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [952, 2048]
NVIDIA B200
23.1µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [8828, 2048]
NVIDIA B200
136.1µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [11006, 2048]
NVIDIA B200
176.3µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [12251, 2048]
NVIDIA B200
182.6µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [12853, 2048]
NVIDIA B200
204.7µs
#2 of 6
2025-10-20
GEMM n5120 k2048fp16 · [14915, 2048]
NVIDIA B200
233.4µs
#1 of 6
2025-10-20
GEMM n5120 k2048fp16 · [16294, 2048]
NVIDIA B200
243.3µs
#1 of 6
2025-10-20

Reported · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:78d3f3c768d3e0404cac6edabacda66636829d892c7ba20fdd57674d0d41e8b3
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16

Kernel source

main.py7 lines
import torch
import torch.nn.functional as F

def run(A: torch.Tensor, B: torch.Tensor):
    C = F.linear(A, B)
    return C

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reported

JSON