Skip to content
KernelIndex
Search⌘K

submission 150025

albanD · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 58 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-150025?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Sortsuite of 5 cases
NVIDIA H100
6.64ms
#17 of 26
2025-12-12

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:4d9a9d803f38e5bfd69c57999f54b6e5f6cd1be945c3b20a23fd31ce0c831b4b
license declaredunknown
license concludedunknown
authorsalbanD
imported2026-08-15

Kernel source

submission.py58 lines
#!POPCORN leaderboard sort_v2

from task import input_t, output_t
import torch
import triton
import triton.language as tl


@triton.jit
def simple_sort_kernel(
    input_ptr,
    output_ptr,
    n_elements,
    BLOCK_SIZE: tl.constexpr,
):
    """Simple sort kernel using tl.sort."""
    offsets = tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements
    data = tl.load(input_ptr + offsets, mask=mask, other=float('inf'))

    # Use Triton's built-in sort
    sorted_data = tl.sort(data)

    tl.store(output_ptr + offsets, sorted_data, mask=mask)


def triton_custom_kernel(data: input_t) -> output_t:
    inp, output = data

    n_elements = inp.numel()

    # Round up to next power of 2 for sort
    block_size = 1
    while block_size < n_elements:
        block_size *= 2

    # Cap block size to reasonable maximum
    block_size = min(block_size, 1024)

    # Launch kernel
    simple_sort_kernel[(1,)](
        inp,
        output,
        n_elements,
        BLOCK_SIZE=block_size,
    )

    return output

def naive_custom_kernel(data):
    inp, out = data

    out.copy_(torch.sort(inp)[0])
    return out


custom_kernel = naive_custom_kernel
scrolls · 58 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 150024.

Best evidence level for this revision: reported

JSON