submission 555130
Narain · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 65 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-gated-deltanet-chunk-fwd-h-555130?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:c3db2fd5d3bf26ac87094b9a3d2af90916d6a252176c6d9290cf3716268dba9f
license declaredunknown
license concludedunknown
authorsNarain
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 4
num_warps=4,stages = 1
num_stages=1,Kernel source
submission.py65 lines
from task import input_t, output_t
import os
os.environ["HELION_AUTOTUNE_EFFORT"] = "none"
import torch
torch.backends.cuda.matmul.allow_tf32 = False
import helion
import helion.language as hl
CHUNK_SIZE = 64
_ACF = "/opt/booster_pack/chunk_fwd_h_1.acf"
@helion.kernel(
dot_precision="ieee",
config=helion.Config(
block_sizes=[16],
num_warps=4,
num_stages=1,
advanced_controls_file=_ACF,
),
)
def gdn_chunk_fwd_h_kernel(
k: torch.Tensor, w: torch.Tensor, u: torch.Tensor, g: torch.Tensor, chunk_size: int,
) -> tuple[torch.Tensor, torch.Tensor]:
B, T, H, K = k.shape; K = hl.specialize(K)
chunk_size = hl.specialize(chunk_size); V = u.shape[-1]
NT = T // chunk_size
h = torch.empty(B, NT, H, K, V, dtype=torch.float32, device=k.device)
v_new = torch.empty_like(u)
block_v = hl.register_block_size(V)
for tile_b, tile_h, tile_v in hl.tile([B, H, V], block_size=[1, 1, block_v]):
i_b = tile_b.id; i_h = tile_h.id
b_h = hl.zeros([K, tile_v], dtype=torch.float32)
for t_i in hl.tile(T, block_size=chunk_size):
h[i_b, t_i.id, i_h, :, tile_v] = b_h
b_w = w[i_b, t_i, i_h, :]
b_v = torch.matmul(b_w, b_h)
p_v = u[i_b, t_i, i_h, tile_v]
b_v = p_v - b_v
v_new[i_b, t_i, i_h, tile_v] = b_v
b_g = g[i_b, t_i, i_h]
t_i_last = min(t_i.begin + chunk_size, T) - 1
b_g_last = g[i_b, t_i_last, i_h]
b_v_gated = hl.inline_triton(
"""
mask = {tidx} < {T}
gate = tl.where(mask, tl.exp({g_last} - {g}), 0.0)
{bv} * gate[:, None]
""",
args={"tidx": t_i.index, "T": T, "g_last": b_g_last, "g": b_g, "bv": b_v},
output_like=b_v,
)
b_h = b_h * torch.exp(b_g_last)
p_k = k[i_b, t_i, i_h, :]
# addmm: fused b_h + k^T @ v_gated
b_h = torch.addmm(b_h, p_k.T, b_v_gated)
return h, v_new
def custom_kernel(data: input_t) -> output_t:
k, w, u, g = data
return gdn_chunk_fwd_h_kernel(k, w, u, g, CHUNK_SIZE)
scrolls · 65 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON