Skip to content
KernelIndex
Search⌘K

submission 99051

Petr_Rocoss · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 101 lines, June 9 Researcher Reciprocity License v1.0.

nvfp4_gemv_ultimate.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-99051?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA H100
376.8ms
#32 of 35
2025-11-23

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:cf83f28eee34f66ee2dc476139f0c46a09bb7d89b6cf438989dc51a6e4991ee1
license declaredunknown
license concludedunknown
authorsPetr_Rocoss
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

autotune@triton.autotune(
num-warps = 8triton.Config({'BLOCK_W': 256}, num_warps=8, num_stages=4),
stages = 4triton.Config({'BLOCK_W': 256}, num_warps=8, num_stages=4),

Kernel source

nvfp4_gemv_ultimate.py101 lines
import torch
import triton
import triton.language as tl

@triton.autotune(
    configs=[
        triton.Config({'BLOCK_W': 256}, num_warps=8, num_stages=4),
        triton.Config({'BLOCK_W': 128}, num_warps=8, num_stages=4),
        triton.Config({'BLOCK_W': 512}, num_warps=8, num_stages=3),
        triton.Config({'BLOCK_W': 64}, num_warps=4, num_stages=4),
    ],
    key=['w_out', 'c_in', 'k_size'],
)
@triton.jit
def conv2d_kernel(
    input_ptr, weight_ptr, output_ptr,
    stride_in_n, stride_in_c, stride_in_h, stride_in_w,
    stride_w_out, stride_w_in, stride_w_h, stride_w_w,
    stride_out_n, stride_out_c, stride_out_h, stride_out_w,
    H_IN, W_IN, H_OUT, W_OUT, C_IN, C_OUT, K,
    BLOCK_W: tl.constexpr
):
    """Оптимизированное ядро Conv2D с правильным порядком циклов."""
    
    pid_w = tl.program_id(0)
    pid_h = tl.program_id(1)
    pid_z = tl.program_id(2)
    
    batch_idx = pid_z // C_OUT
    out_ch = pid_z % C_OUT
    out_h = pid_h
    
    # Width offsets
    offs_w = pid_w * BLOCK_W + tl.arange(0, BLOCK_W)
    mask_w = offs_w < W_OUT
    
    # Base pointers
    out_base = output_ptr + batch_idx * stride_out_n + out_ch * stride_out_c + out_h * stride_out_h
    in_base = input_ptr + batch_idx * stride_in_n + out_h * stride_in_h
    wei_base = weight_ptr + out_ch * stride_w_out
    
    acc = tl.zeros([BLOCK_W], dtype=tl.float32)
    
    # === ОПТИМИЗИРОВАННЫЙ ПОРЯДОК ЦИКЛОВ ===
    # cin (внешний) -> kh, kw (внутренние) для лучшей локальности памяти
    for cin in range(C_IN):
        in_ch_ptr = in_base + cin * stride_in_c
        wei_ch_ptr = wei_base + cin * stride_w_in
        
        for kh in range(K):
            in_row_ptr = in_ch_ptr + kh * stride_in_h
            wei_row_ptr = wei_ch_ptr + kh * stride_w_h
            
            for kw in range(K):
                # Загрузка веса - минимальный overhead
                wei_val = tl.load(wei_row_ptr + kw * stride_w_w)
                
                # Загрузка входа - правильные индексы
                # Input: [batch, cin, out_h+kh, out_w+kw]
                in_ptrs = in_row_ptr + (offs_w + kw) * stride_in_w
                in_val = tl.load(in_ptrs, mask=mask_w, other=0.0)
                
                # Аккумуляция
                acc = acc + in_val * wei_val
    
    # Запись
    tl.store(out_base + offs_w * stride_out_w, acc, mask=mask_w)


def custom_kernel(data):
    """Вход: (input_tensor, kernel, output_tensor)."""
    
    input_tensor, kernel, output_tensor = data
    
    input_tensor = input_tensor.contiguous()
    kernel = kernel.contiguous()
    
    batch, c_in, h_in, w_in = input_tensor.shape
    c_out, _, k_h, k_w = kernel.shape
    
    h_out = h_in - k_h + 1
    w_out = w_in - k_w + 1
    
    grid = lambda META: (
        triton.cdiv(w_out, META['BLOCK_W']),
        h_out,
        batch * c_out
    )
    
    conv2d_kernel[grid](
        input_tensor, kernel, output_tensor,
        *input_tensor.stride(),
        *kernel.stride(),
        *output_tensor.stride(),
        h_in, w_in, h_out, w_out,
        c_in, c_out, k_h,
    )
    
    return output_tensor

scrolls · 101 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 99036.

Best evidence level for this revision: reported

JSON