Skip to content
KernelIndex
Search⌘K

submission 555130

Narain · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 65 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-gated-deltanet-chunk-fwd-h-555130?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVIDIA B200
18.6µs
#15 of 28
2026-03-15

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:c3db2fd5d3bf26ac87094b9a3d2af90916d6a252176c6d9290cf3716268dba9f
license declaredunknown
license concludedunknown
authorsNarain
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 4num_warps=4,
stages = 1num_stages=1,

Kernel source

submission.py65 lines
from task import input_t, output_t

import os
os.environ["HELION_AUTOTUNE_EFFORT"] = "none"

import torch
torch.backends.cuda.matmul.allow_tf32 = False

import helion
import helion.language as hl

CHUNK_SIZE = 64
_ACF = "/opt/booster_pack/chunk_fwd_h_1.acf"


@helion.kernel(
    dot_precision="ieee",
    config=helion.Config(
        block_sizes=[16],
        num_warps=4,
        num_stages=1,
        advanced_controls_file=_ACF,
    ),
)
def gdn_chunk_fwd_h_kernel(
    k: torch.Tensor, w: torch.Tensor, u: torch.Tensor, g: torch.Tensor, chunk_size: int,
) -> tuple[torch.Tensor, torch.Tensor]:
    B, T, H, K = k.shape; K = hl.specialize(K)
    chunk_size = hl.specialize(chunk_size); V = u.shape[-1]
    NT = T // chunk_size
    h = torch.empty(B, NT, H, K, V, dtype=torch.float32, device=k.device)
    v_new = torch.empty_like(u)
    block_v = hl.register_block_size(V)
    for tile_b, tile_h, tile_v in hl.tile([B, H, V], block_size=[1, 1, block_v]):
        i_b = tile_b.id; i_h = tile_h.id
        b_h = hl.zeros([K, tile_v], dtype=torch.float32)
        for t_i in hl.tile(T, block_size=chunk_size):
            h[i_b, t_i.id, i_h, :, tile_v] = b_h
            b_w = w[i_b, t_i, i_h, :]
            b_v = torch.matmul(b_w, b_h)
            p_v = u[i_b, t_i, i_h, tile_v]
            b_v = p_v - b_v
            v_new[i_b, t_i, i_h, tile_v] = b_v
            b_g = g[i_b, t_i, i_h]
            t_i_last = min(t_i.begin + chunk_size, T) - 1
            b_g_last = g[i_b, t_i_last, i_h]
            b_v_gated = hl.inline_triton(
                """
                mask = {tidx} < {T}
                gate = tl.where(mask, tl.exp({g_last} - {g}), 0.0)
                {bv} * gate[:, None]
                """,
                args={"tidx": t_i.index, "T": T, "g_last": b_g_last, "g": b_g, "bv": b_v},
                output_like=b_v,
            )
            b_h = b_h * torch.exp(b_g_last)
            p_k = k[i_b, t_i, i_h, :]
            # addmm: fused b_h + k^T @ v_gated
            b_h = torch.addmm(b_h, p_k.T, b_v_gated)
    return h, v_new

def custom_kernel(data: input_t) -> output_t:
    k, w, u, g = data
    return gdn_chunk_fwd_h_kernel(k, w, u, g, CHUNK_SIZE)
scrolls · 65 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON