submission 880494
HappyAny · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 35 lines, June 9 Researcher Reciprocity License v1.0.
main.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-cholesky-880494?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:e2ac82a34d6501c91d498084703d10ac3a75a0ba1605f4a0f6e555d9260504a6
license declaredunknown
license concludedunknown
authorsHappyAny
imported2026-08-26
Kernel source
main.py35 lines
import torch
from task import input_t, output_t
def _fallback(data: torch.Tensor) -> torch.Tensor:
return torch.linalg.cholesky_ex(data, check_errors=False).L
def _loop_batch(data: torch.Tensor) -> torch.Tensor:
out = torch.empty_like(data)
for i in range(data.shape[0]):
out[i : i + 1].copy_(
torch.linalg.cholesky_ex(data[i : i + 1], check_errors=False).L
)
return out
def custom_kernel(data: input_t) -> output_t:
n = data.shape[-1]
batch = data.shape[0]
# Known win from loop_4096_batch2_experimental.py.
if n == 4096 and batch == 2:
return _loop_batch(data)
# Probe whether other low-batch large shapes have the same bad batched path.
if n == 2048 and batch in (2, 8):
return _loop_batch(data)
if n == 1024 and batch == 4:
return _loop_batch(data)
return _fallback(data)
scrolls · 35 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON