Skip to content
KernelIndex
Search⌘K

submission 780441

Kernel-Zhang · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 28 lines, June 9 Researcher Reciprocity License v1.0.

ref.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-780441?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA A100
635.2µs
#4 of 27
2026-04-28

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:b83f48e6ea4749605c7a1480ba7559f33d9a6d79e3a9712b2009cff7db2aef0e
license declaredunknown
license concludedunknown
authorsKernel-Zhang
imported2026-08-15

Kernel source

ref.py28 lines
import torch
from task import input_t, output_t
from utils import make_match_reference, DeterministicContext

def custom_kernel(data: input_t) -> output_t:
    a, b, c = data
    return a @ b

def generate_input(m: int, n: int, k: int, seed: int) -> input_t:
    gen = torch.Generator(device='cuda')
    gen.manual_seed(seed)
    a = torch.empty(m, k, device='cuda', dtype=torch.float16)
    a.uniform_(0, 1, generator=gen)
    b = torch.empty(k, n, device='cuda', dtype=torch.float16)
    b.uniform_(0, 1, generator=gen)
    c = torch.empty(m, n, device='cuda', dtype=torch.float16)
    return a, b, c


def ref_kernel(data: input_t) -> output_t:
    with DeterministicContext():
        a, b, c = data
        return a @ b


check_implementation = make_match_reference(ref_kernel)

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON