Skip to content
KernelIndex
Search⌘K

submission 870564

Ayushi Nanavati · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 35 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-eigh-870564?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVIDIA B200
48.8ms
#152 of 286
2026-07-12

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:1424e63b5ab94e32b13d1c3c05557361516a4320370fb62787a91accff219343
license declaredunknown
license concludedunknown
authorsAyushi Nanavati
imported2026-08-26

Kernel source

submission.py35 lines
"""GPU MODE leaderboard 775 - batched real symmetric eigh on B200.

Submission 1: clean dense baseline.

Recon (see RECON.md) established that EVERY timed benchmark case is a full
dense matrix. There are no diagonal / identity / zero / band-sparse cases in
the benchmark set (the maintainers explicitly removed the one diagonal
benchmark row, and a Triton diagonal fast-path measured *slower* on benchmark:
45.9s vs 28.0s baseline, because structure detection costs on every dense
case for zero benefit). Therefore a structure-dispatch submission cannot move
the geomean and would only add per-batch reduction overhead on the dense path.

So submission 1 is the plain, correct baseline. It exists to:
  * validate the popcorn toolchain end-to-end on the runner, and
  * record our own scored baseline number to compare future work against.

The real lever (beating baseline) requires low-precision / tensor-core spectral
methods -- out of scope for submission 1, see RECON.md "Next step".

Contract: custom_kernel(A: batch x n x n float32 CUDA) -> (Q, L) with
  Q: batch x n x n float32, columns are orthonormal eigenvectors
  L: batch x n float32, eigenvalues ascending
matching torch.linalg.eigh(A) == (L, Q)  ->  we return (Q, L).
"""

import torch
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    # torch.linalg.eigh returns (eigenvalues ascending, eigenvectors as columns).
    # The task convention wants (Q, L), so we swap the order.
    values, vectors = torch.linalg.eigh(data)
    return vectors, values
scrolls · 35 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON