Skip to content
KernelIndex
Search⌘K

submission 838388

jashanpreetsingh1 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 20 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-eigh-838388?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVIDIA B200
53.8ms
#213 of 286
2026-06-27

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:bea337b42c07bfb403dfa3f9aeab2fbbf428c83d43aa168ca2ea57080d175b05
license declaredunknown
license concludedunknown
authorsjashanpreetsingh1
imported2026-08-26

Kernel source

submission.py20 lines
#!POPCORN leaderboard eigh
#!POPCORN gpu B200

# Batched real-symmetric eigendecomposition.
#
# Baseline: torch.linalg.eigh dispatches to cuSOLVER's batched Jacobi (syevj),
# which is well-optimized at every size we hit (B200 profiling confirmed it uses
# batch_parallel_jacobi for n=32 and syevj for n=512). A naive custom Jacobi was
# measured ~5x slower at n=32, so there is no cheap win in the small cases.
# Beating this requires a fundamentally faster algorithm (e.g. a low-precision
# tensor-core spectral solver exploiting the loose residual gates).

import torch
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    values, vectors = torch.linalg.eigh(data)
    return vectors, values

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON