Skip to content
KernelIndex
Search⌘K

torch_matmul_317103

FlashInfer-Bench baselines · python · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 7 lines, Apache-2.0, pinned at da91508.

main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-torch-matmul-317103?include=source"
interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
measured onNVIDIA B200
declared hardwareCPU, NVIDIA A100, NVIDIA B200, NVIDIA H100
architecturessm_100, sm_90, unknown
dtypesfp16

Benchmark evidence

25 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n128 k2048fp16 · [34, 2048]
NVIDIA B200
9.45µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [25, 2048]
NVIDIA B200
9.48µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [128, 2048]
NVIDIA B200
9.51µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [63, 2048]
NVIDIA B200
9.57µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [6, 2048]
NVIDIA B200
9.63µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [17, 2048]
NVIDIA B200
9.71µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [2, 2048]
NVIDIA B200
9.84µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [4, 2048]
NVIDIA B200
9.89µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [64, 2048]
NVIDIA B200
9.94µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [8, 2048]
NVIDIA B200
9.99µs
#1 of 7
2025-10-20
Show all 25 measurements ›
GEMM n128 k2048fp16 · [32, 2048]
NVIDIA B200
10.0µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [1, 2048]
NVIDIA B200
10.1µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [93, 2048]
NVIDIA B200
10.1µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [5, 2048]
NVIDIA B200
10.2µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [172, 2048]
NVIDIA B200
10.2µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [289, 2048]
NVIDIA B200
10.2µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [492, 2048]
NVIDIA B200
10.3µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [16, 2048]
NVIDIA B200
10.4µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [952, 2048]
NVIDIA B200
10.7µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [8828, 2048]
NVIDIA B200
14.8µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [11006, 2048]
NVIDIA B200
16.7µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [12251, 2048]
NVIDIA B200
18.5µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [12853, 2048]
NVIDIA B200
18.9µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [14915, 2048]
NVIDIA B200
21.0µs
#1 of 7
2025-10-20
GEMM n128 k2048fp16 · [16294, 2048]
NVIDIA B200
22.5µs
#1 of 7
2025-10-20

Reported · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:ae14a40a97638f12f4f1d61def04d5bea044bd4d5a5ee13ab3f3af5c2efa04d6
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16

Kernel source

main.py7 lines
import torch
import torch.nn.functional as F

def run(A: torch.Tensor, B: torch.Tensor):
    C = F.linear(A, B)
    return C

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reported

JSON