Skip to content
KernelIndex
Search⌘K

submission 553189

brandonin · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 70 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-fp8-quant-553189?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVIDIA B200
14.8µs
#7 of 17
2026-03-14

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:eed19ef602b1476ad75129daf2b4e151db76909c5b1cdcc91e2d22de29148d17
license declaredunknown
license concludedunknown
authorsbrandonin
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 4(1, 256, 64): helion.Config(block_sizes=[4], num_stages=1, num_warps=4, pid_type='flat'),
stages = 1(1, 256, 64): helion.Config(block_sizes=[4], num_stages=1, num_warps=4, pid_type='flat'),

Kernel source

submission.py70 lines
from task import input_t, output_t

import torch
import helion
import helion.language as hl


# Per-shape configs: B200-autotuned with HELION_AUTOTUNE_EFFORT=full
SHAPE_CONFIGS: dict[tuple, helion.Config] = {
    # Test shapes (autotuned on B200)
    (1, 256, 64): helion.Config(block_sizes=[4], num_stages=1, num_warps=4, pid_type='flat'),
    (4, 512, 128): helion.Config(block_sizes=[16], num_stages=1, num_warps=4, pid_type='flat'),
    (16, 1024, 64): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
    (1, 4096, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
    (8, 4096, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
    # Benchmark shapes (autotuned on B200)
    (16, 4096, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
    (256, 4096, 128): helion.Config(block_sizes=[8], indexing=['pointer', 'pointer', 'pointer', 'pointer', 'pointer', 'tensor_descriptor'], load_eviction_policies=['', '', 'first'], num_stages=2, num_warps=4, pid_type='flat'),
    (256, 8192, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
    (4096, 7168, 128): helion.Config(block_sizes=[8], indexing=['pointer', 'pointer', 'pointer', 'pointer', 'tensor_descriptor', 'pointer'], load_eviction_policies=['first', '', ''], num_stages=5, num_warps=4, pid_type='flat'),
}


def _make_kernel(config: helion.Config):
    @helion.kernel(static_shapes=True, config=config)
    def kernel(
        data: torch.Tensor,        # [N, G] input rows
        scales_out: torch.Tensor,  # [N] output normalization factors
    ) -> torch.Tensor:
        nrows = data.size(0)
        ncols = hl.specialize(data.size(1))
        MAX_VAL = 448.0
        EPS = 1e-10

        qout = torch.empty(nrows, ncols, dtype=torch.float32, device=data.device)

        for rr in hl.tile(nrows):
            row = data[rr, :].to(torch.float32)
            amax = torch.amax(torch.abs(row), -1)
            amax = torch.clamp(amax, min=EPS)
            scale = amax / MAX_VAL
            qout[rr, :] = torch.clamp(row / scale[:, None], -MAX_VAL, MAX_VAL)
            scales_out[rr] = scale

        return qout

    return kernel


_KERNELS = {shape: _make_kernel(cfg) for shape, cfg in SHAPE_CONFIGS.items()}


def custom_kernel(data: input_t) -> output_t:
    x, x_q, x_s = data
    T, H = x.shape
    G = x_s.shape[1]
    gsz = H // G
    N = T * G

    kernel = _KERNELS[(T, H, gsz)]

    flat_in = x.reshape(N, gsz)
    flat_s = x_s.reshape(N)

    flat_q = kernel(flat_in, flat_s)

    x_q[...] = flat_q.reshape(T, H)
    x_s[...] = flat_s.reshape(T, G)
    return x_q, x_s
scrolls · 70 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 552685.

- #!POPCORN leaderboard fp8_quant
- #!POPCORN gpu B200_Nebius
-
from task import input_t, output_t
import torch
⋯ 1 unchanged lines
import helion.language as hl
- # Per-shape configs: map (num_tokens, hidden_dim, group_size) to optimized helion.Config objects.
+ # Per-shape configs: B200-autotuned with HELION_AUTOTUNE_EFFORT=full
SHAPE_CONFIGS: dict[tuple, helion.Config] = {
- # Test shapes
- (1, 256, 64): helion.Config(block_sizes=[4], num_warps=4, num_stages=2),
- (4, 512, 128): helion.Config(block_sizes=[16], num_warps=4, num_stages=2),
- (16, 1024, 64): helion.Config(block_sizes=[64], num_warps=4, num_stages=2),
- (1, 4096, 128): helion.Config(block_sizes=[32], num_warps=4, num_stages=2),
- (8, 4096, 128): helion.Config(block_sizes=[64], num_warps=4, num_stages=2),
- # Benchmark shapes
- (16, 4096, 128): helion.Config(block_sizes=[64], num_warps=4, num_stages=2),
- (256, 4096, 128): helion.Config(block_sizes=[128], num_warps=4, num_stages=2),
- (256, 8192, 128): helion.Config(block_sizes=[128], num_warps=4, num_stages=2),
- (4096, 7168, 128): helion.Config(block_sizes=[128], num_warps=4, num_stages=2),
+ # Test shapes (autotuned on B200)
+ (1, 256, 64): helion.Config(block_sizes=[4], num_stages=1, num_warps=4, pid_type='flat'),
+ (4, 512, 128): helion.Config(block_sizes=[16], num_stages=1, num_warps=4, pid_type='flat'),
+ (16, 1024, 64): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
+ (1, 4096, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
+ (8, 4096, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
+ # Benchmark shapes (autotuned on B200)
+ (16, 4096, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
+ (256, 4096, 128): helion.Config(block_sizes=[8], indexing=['pointer', 'pointer', 'pointer', 'pointer', 'pointer', 'tensor_descriptor'], load_eviction_policies=['', '', 'first'], num_stages=2, num_warps=4, pid_type='flat'),
+ (256, 8192, 128): helion.Config(block_sizes=[32], num_stages=1, num_warps=4, pid_type='flat'),
+ (4096, 7168, 128): helion.Config(block_sizes=[8], indexing=['pointer', 'pointer', 'pointer', 'pointer', 'tensor_descriptor', 'pointer'], load_eviction_policies=['first', '', ''], num_stages=5, num_warps=4, pid_type='flat'),
}
scrolls · 38 diff lines total

Best evidence level for this revision: reported

JSON