Skip to content
KernelIndex
Search⌘K

submission 880494

HappyAny · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 35 lines, June 9 Researcher Reciprocity License v1.0.

main.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-cholesky-880494?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVIDIA B200
1.82ms
#234 of 337
2026-07-16

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:e2ac82a34d6501c91d498084703d10ac3a75a0ba1605f4a0f6e555d9260504a6
license declaredunknown
license concludedunknown
authorsHappyAny
imported2026-08-26

Kernel source

main.py35 lines
import torch

from task import input_t, output_t


def _fallback(data: torch.Tensor) -> torch.Tensor:
    return torch.linalg.cholesky_ex(data, check_errors=False).L


def _loop_batch(data: torch.Tensor) -> torch.Tensor:
    out = torch.empty_like(data)
    for i in range(data.shape[0]):
        out[i : i + 1].copy_(
            torch.linalg.cholesky_ex(data[i : i + 1], check_errors=False).L
        )
    return out


def custom_kernel(data: input_t) -> output_t:
    n = data.shape[-1]
    batch = data.shape[0]

    # Known win from loop_4096_batch2_experimental.py.
    if n == 4096 and batch == 2:
        return _loop_batch(data)

    # Probe whether other low-batch large shapes have the same bad batched path.
    if n == 2048 and batch in (2, 8):
        return _loop_batch(data)

    if n == 1024 and batch == 4:
        return _loop_batch(data)

    return _fallback(data)
scrolls · 35 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON