Skip to content
KernelIndex
Search⌘K

submission 763415

CaptnJackSparrow · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 48 lines, June 9 Researcher Reciprocity License v1.0.

submission_cuda_inline_H100.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-763415?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Sortsuite of 5 cases
NVIDIA H100
2.04ms
#3 of 26
2026-04-12

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:1111798e5f9518a0b46b2f665150322cdc8fd614fb3fa367bb14e19addad53e4
license declaredunknown
license concludedunknown
authorsCaptnJackSparrow
imported2026-08-15

Kernel source

submission_cuda_inline_H100.py48 lines
import torch
from torch.utils.cpp_extension import load_inline
from task import input_t, output_t

sort_cuda_source = """
#include <cub/cub.cuh>
#include <cuda_runtime.h>

static void* d_temp = nullptr;
static size_t temp_bytes = 0;

void sort_cuda(torch::Tensor input, torch::Tensor output) {
    int N = input.numel();
    float* d_in = input.data_ptr<float>();
    float* d_out = output.data_ptr<float>();

    size_t required = 0;
    cub::DeviceRadixSort::SortKeys(nullptr, required, d_in, d_out, N);

    if (required > temp_bytes) {
        if (d_temp) cudaFree(d_temp);
        cudaMalloc(&d_temp, required);
        temp_bytes = required;
    }

    cub::DeviceRadixSort::SortKeys(d_temp, temp_bytes, d_in, d_out, N);
}
"""

sort_cpp_source = """
#include <torch/extension.h>
void sort_cuda(torch::Tensor input, torch::Tensor output);
"""

sort_module = load_inline(
    name='sort_cuda',
    cpp_sources=sort_cpp_source,
    cuda_sources=sort_cuda_source,
    functions=['sort_cuda'],
    verbose=True,
    extra_cuda_cflags=['-O3', '--use_fast_math', '-gencode', 'arch=compute_90,code=sm_90'],
)

def custom_kernel(data: input_t) -> output_t:
    data, output = data
    sort_module.sort_cuda(data, output)
    return output
scrolls · 48 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON