Skip to content
KernelIndex
Search⌘K

submission 409172

novo_force · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 64 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-trimul-409172?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVIDIA B200
7.24ms
#33 of 43
2026-01-29

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:a80a08f2c668054743b0f1b93d62400fbfd11ba94cf9d2ad8b31c881d6498606
license declaredunknown
license concludedunknown
authorsnovo_force
imported2026-08-15

Kernel source

submission.py64 lines
from __future__ import annotations

from typing import Dict, Tuple, Any

import torch
import torch.nn.functional as F


def _outgoing_core(left: torch.Tensor, right: torch.Tensor) -> torch.Tensor:
    
    
    
    bs, i, k, hidden = left.shape
    j = right.shape[1]

    left_bd = left.permute(0, 3, 1, 2).contiguous().view(bs * hidden, i, k)
    right_bd = right.permute(0, 3, 1, 2).contiguous().view(bs * hidden, j, k)
    out_bd = torch.bmm(left_bd, right_bd.transpose(1, 2))
    return out_bd.view(bs, hidden, i, j).permute(0, 2, 3, 1).contiguous()


@torch.inference_mode()
def custom_kernel(data: Tuple[torch.Tensor, torch.Tensor, Dict[str, torch.Tensor], Dict[str, Any]]) -> torch.Tensor:
    x, mask, weights, config = data

    dim = int(config["dim"])
    hidden_dim = int(config["hidden_dim"])

    if x.dtype != torch.float32:
        x = x.to(dtype=torch.float32)

    x = F.layer_norm(x, (dim,), weights["norm.weight"], weights["norm.bias"], 1e-5)

    left = F.linear(x, weights["left_proj.weight"], None)
    right = F.linear(x, weights["right_proj.weight"], None)

    mask_f = mask.unsqueeze(-1)
    if mask_f.dtype != left.dtype:
        mask_f = mask_f.to(dtype=left.dtype)
    left = left * mask_f
    right = right * mask_f

    left_gate = torch.sigmoid(F.linear(x, weights["left_gate.weight"], None))
    right_gate = torch.sigmoid(F.linear(x, weights["right_gate.weight"], None))
    out_gate = torch.sigmoid(F.linear(x, weights["out_gate.weight"], None))

    left = left * left_gate
    right = right * right_gate

    out = _outgoing_core(left, right)
    out = F.layer_norm(
        out,
        (hidden_dim,),
        weights["to_out_norm.weight"],
        weights["to_out_norm.bias"],
        1e-5,
    )
    out = out * out_gate
    out = F.linear(out, weights["to_out.weight"], None)
    return out


__all__ = ["custom_kernel"]
scrolls · 64 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON