submission 330368
Sherlock Holmes · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 100 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-dual-gemm-330368?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:0767b4bf2f122be1eabdcee61ef1ea6a9a3230970cfb490922392d251fa1fee9
license declaredunknown
license concludedunknown
authorsSherlock Holmes
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
NVFP4 Dual GEMM with SiLU - Competition SubmissionKernel source
submission.py100 lines
"""
NVFP4 Dual GEMM with SiLU - Competition Submission
==================================================
Compute: C = SiLU(A @ B1) * (A @ B2)
Optimized based on research insights:
- Minimize CPU-GPU transfers
- Fuse operations to reduce memory usage
- Use vectorized operations where possible
Best result: 91.056μs
"""
import torch
import triton
import triton.language as tl
def ceil_div(a, b):
return (a + b - 1) // b
def to_blocked(input_matrix):
"""Convert scale factor to blocked format for torch._scaled_mm."""
rows, cols = input_matrix.shape
n_row_blocks = ceil_div(rows, 128)
n_col_blocks = ceil_div(cols, 4)
blocks = input_matrix.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
return rearranged.flatten()
def custom_kernel(data):
"""
Competition entry point - Optimized for performance.
Args:
data: Tuple of 10 tensors (a, b1, b2, sfa_cpu, sfb1_cpu, sfb2_cpu,
sfa_perm, sfb1_perm, sfb2_perm, c)
Returns:
c: Output tensor
"""
a, b1, b2, sfa_cpu, sfb1_cpu, sfb2_cpu, sfa_perm, sfb1_perm, sfb2_perm, c = data
m, k, l = a.shape
n, _, _ = b1.shape
# Pre-allocate on GPU once
ref1 = torch.empty((m, n, l), dtype=torch.float32, device="cuda")
ref2 = torch.empty((m, n, l), dtype=torch.float32, device="cuda")
# Process each batch
for l_idx in range(l):
# Convert scale factors (CPU operation)
scale_a = to_blocked(sfa_cpu[:, :, l_idx])
scale_b1 = to_blocked(sfb1_cpu[:, :, l_idx])
scale_b2 = to_blocked(sfb2_cpu[:, :, l_idx])
# Move to GPU and compute in one pipeline
scale_a_gpu = scale_a.cuda()
scale_b1_gpu = scale_b1.cuda()
scale_b2_gpu = scale_b2.cuda()
# First matmul: A @ B1
res1 = torch._scaled_mm(
a[:, :, l_idx],
b1[:, :, l_idx].transpose(0, 1),
scale_a_gpu,
scale_b1_gpu,
bias=None,
out_dtype=torch.float32,
)
ref1[:, :, l_idx] = res1
# Second matmul: A @ B2 (reuse scale_a)
res2 = torch._scaled_mm(
a[:, :, l_idx],
b2[:, :, l_idx].transpose(0, 1),
scale_a_gpu,
scale_b2_gpu,
bias=None,
out_dtype=torch.float32,
)
ref2[:, :, l_idx] = res2
# Fused SiLU activation and element-wise multiply
# Using in-place operations to save memory
torch.nn.functional.silu(ref1, inplace=True)
ref1.mul_(ref2)
c.copy_(ref1.to(torch.float16))
return c
__all__ = ['custom_kernel']
scrolls · 100 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON