Skip to content
KernelIndex
Search⌘K

torch / matmul67278e

torch_matmul_67278e · FlashInfer-Bench baselines · python · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 7 lines, Apache-2.0, pinned at da91508.

main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-torch-matmul-67278e?include=source"
interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
measured onNVIDIA B200
declared hardwareCPU, NVIDIA A100, NVIDIA B200, NVIDIA H100
architecturessm_100, sm_90, unknown
dtypesfp16

Benchmark evidence

17 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n256 k7168fp16 · [16, 7168]
NVIDIA B200
15.7µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [80, 7168]
NVIDIA B200
15.7µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [63, 7168]
NVIDIA B200
15.7µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [15, 7168]
NVIDIA B200
16.0µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [57, 7168]
NVIDIA B200
16.1µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [56, 7168]
NVIDIA B200
16.1µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [4, 7168]
NVIDIA B200
16.1µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [54, 7168]
NVIDIA B200
16.1µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [58, 7168]
NVIDIA B200
16.1µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [14, 7168]
NVIDIA B200
16.2µs
#1 of 7
2025-10-20
Show all 17 measurements ›
GEMM n256 k7168fp16 · [55, 7168]
NVIDIA B200
16.3µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [32, 7168]
NVIDIA B200
16.3µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [1, 7168]
NVIDIA B200
16.3µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [53, 7168]
NVIDIA B200
16.3µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [901, 7168]
NVIDIA B200
18.7µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [11948, 7168]
NVIDIA B200
49.1µs
#1 of 7
2025-10-20
GEMM n256 k7168fp16 · [14104, 7168]
NVIDIA B200
53.4µs
#1 of 7
2025-10-20

Reported · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:d7e2359135dd88d96fd7425d96e9f12d6ba8727a7aebc0d71b3d6c1b16deeafc
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16

Kernel source

main.py7 lines
import torch
import torch.nn.functional as F

def run(A: torch.Tensor, B: torch.Tensor):
    C = F.linear(A, B)
    return C

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reported

JSON