Skip to content
KernelIndex
Search⌘K

torch / matmul926adc

torch_matmul_926adc · FlashInfer-Bench baselines · python · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 7 lines, Apache-2.0, pinned at da91508.

main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-torch-matmul-926adc?include=source"
interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
measured onNVIDIA B200
declared hardwareCPU, NVIDIA A100, NVIDIA B200, NVIDIA H100
architecturessm_100, sm_90, unknown
dtypesfp16

Benchmark evidence

29 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n2048 k4096fp16 · [17, 4096]
NVIDIA B200
13.8µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [34, 4096]
NVIDIA B200
14.0µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [16, 4096]
NVIDIA B200
14.1µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [25, 4096]
NVIDIA B200
14.1µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [32, 4096]
NVIDIA B200
14.2µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [64, 4096]
NVIDIA B200
14.3µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [93, 4096]
NVIDIA B200
14.3µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [15, 4096]
NVIDIA B200
14.3µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [128, 4096]
NVIDIA B200
14.3µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [63, 4096]
NVIDIA B200
14.4µs
#1 of 7
2025-10-20
Show all 29 measurements ›
GEMM n2048 k4096fp16 · [172, 4096]
NVIDIA B200
14.4µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [8, 4096]
NVIDIA B200
14.7µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [289, 4096]
NVIDIA B200
14.7µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [2, 4096]
NVIDIA B200
15.8µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [5, 4096]
NVIDIA B200
16.2µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [6, 4096]
NVIDIA B200
16.3µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [4, 4096]
NVIDIA B200
16.6µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [492, 4096]
NVIDIA B200
16.9µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [1, 4096]
NVIDIA B200
17.2µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [969, 4096]
NVIDIA B200
20.6µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [952, 4096]
NVIDIA B200
22.8µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [8828, 4096]
NVIDIA B200
99.0µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [11938, 4096]
NVIDIA B200
135.3µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [11006, 4096]
NVIDIA B200
142.9µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [12853, 4096]
NVIDIA B200
146.8µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [12251, 4096]
NVIDIA B200
153.0µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [14915, 4096]
NVIDIA B200
170.6µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [15813, 4096]
NVIDIA B200
201.5µs
#1 of 7
2025-10-20
GEMM n2048 k4096fp16 · [16294, 4096]
NVIDIA B200
209.8µs
#1 of 7
2025-10-20

Reported · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:8e3fd7d54fe1282f5458444fbf9b7320c83325b2da9023dbf1bc0f10d37ac85a
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16

Kernel source

main.py7 lines
import torch
import torch.nn.functional as F

def run(A: torch.Tensor, B: torch.Tensor):
    C = F.linear(A, B)
    return C

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reported

JSON