submission 838388
jashanpreetsingh1 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 20 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-eigh-838388?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:bea337b42c07bfb403dfa3f9aeab2fbbf428c83d43aa168ca2ea57080d175b05
license declaredunknown
license concludedunknown
authorsjashanpreetsingh1
imported2026-08-26
Kernel source
submission.py20 lines
#!POPCORN leaderboard eigh
#!POPCORN gpu B200
# Batched real-symmetric eigendecomposition.
#
# Baseline: torch.linalg.eigh dispatches to cuSOLVER's batched Jacobi (syevj),
# which is well-optimized at every size we hit (B200 profiling confirmed it uses
# batch_parallel_jacobi for n=32 and syevj for n=512). A naive custom Jacobi was
# measured ~5x slower at n=32, so there is no cheap win in the small cases.
# Beating this requires a fundamentally faster algorithm (e.g. a low-precision
# tensor-core spectral solver exploiting the loose residual gates).
import torch
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
values, vectors = torch.linalg.eigh(data)
return vectors, values
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON