Skip to content
KernelIndex
Search⌘K

submission 132573

Nader · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 48 lines, June 9 Researcher Reciprocity License v1.0.

sort_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-132573?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Sortsuite of 5 cases
NVIDIA B200
2.27ms
#12 of 23
2025-12-08

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:71d42569d255131173c892dacda33f91cce4f1d35355e6ceaed63d3804600b1f
license declaredunknown
license concludedunknown
authorsNader
imported2026-08-15

Kernel source

sort_v2.py48 lines
#!POPCORN leaderboard sort_v2

# This is a submission template for popcorn leaderboard 'sort_v2'.
# Your task is as follows:
# > Implement a sort kernel that matches the reference implementation.
# > The kernel should sort the input array in ascending order using a sort algorithm of your choice.
# > 
# > Input arrays are generated as random floating-point numbers, where each row of a roughly square matrix
# > is drawn from a normal distribution with a different mean value per row based on the seed and then flattened into a 1D array.
# The deadline for this leaderboard is 2025-12-30 00:00:00+00:00

# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!

import subprocess

subprocess.run(["pip", "install", "cuda-cccl[cu12]==0.4.1"])

from task import input_t, output_t

import cuda.compute
from cuda.compute import (
    DoubleBuffer,
    SortOrder,
)
import torch

build_in_keys = torch.empty(1, dtype=torch.float32, device="cuda")
build_out_keys = torch.empty(1, dtype=torch.float32, device="cuda")

build_double_buffer = DoubleBuffer(build_in_keys, build_out_keys)

sorter = cuda.compute.make_radix_sort(build_double_buffer, None, None, None, SortOrder.ASCENDING)

input_size = 100000000

temp_storage_size = sorter(None, build_double_buffer, None, None, None, input_size)
d_temp_storage = torch.empty(temp_storage_size, dtype=torch.uint8, device="cuda")

def custom_kernel(data: input_t) -> output_t:
    d_in_keys, d_out_keys = data
    double_buffer = DoubleBuffer(d_in_keys, d_out_keys)
    size = len(d_in_keys)

    sorter(d_temp_storage, double_buffer, None, None, None, size)

    return double_buffer.current()
scrolls · 48 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON