Skip to content
KernelIndex
Search⌘K

claude-opus-4-1 / triton48d048

claude-opus-4-1_triton_48d048 · claude-opus-4-1-20250805 · triton · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 115 lines, Apache-2.0, pinned at da91508.

main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-claude-opus-4-1-triton-48d048?include=source"
interfacetriton
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

43 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n6144 k4096fp16 · [1, 4096]
NVIDIA B200
70.4µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [48, 4096]
NVIDIA B200
70.4µs
#4 of 5
2025-10-16
GEMM n6144 k4096fp16 · [8, 4096]
NVIDIA B200
70.6µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [7, 4096]
NVIDIA B200
70.7µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [32, 4096]
NVIDIA B200
70.7µs
#4 of 5
2025-10-16
GEMM n6144 k4096fp16 · [15, 4096]
NVIDIA B200
70.7µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [16, 4096]
NVIDIA B200
70.8µs
#3 of 5
2025-10-16
GEMM n6144 k4096fp16 · [35, 4096]
NVIDIA B200
70.9µs
#4 of 5
2025-10-16
GEMM n6144 k4096fp16 · [4, 4096]
NVIDIA B200
70.9µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [72, 4096]
NVIDIA B200
71.0µs
#3 of 5
2025-10-16
Show all 43 measurements ›
GEMM n6144 k4096fp16 · [2, 4096]
NVIDIA B200
71.0µs
#3 of 5
2025-10-16
GEMM n6144 k4096fp16 · [24, 4096]
NVIDIA B200
71.1µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [56, 4096]
NVIDIA B200
71.2µs
#4 of 5
2025-10-16
GEMM n6144 k4096fp16 · [70, 4096]
NVIDIA B200
71.2µs
#4 of 5
2025-10-16
GEMM n6144 k4096fp16 · [64, 4096]
NVIDIA B200
71.3µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [40, 4096]
NVIDIA B200
71.4µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [80, 4096]
NVIDIA B200
71.6µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [144, 4096]
NVIDIA B200
71.7µs
#3 of 5
2025-10-16
GEMM n6144 k4096fp16 · [136, 4096]
NVIDIA B200
71.7µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [160, 4096]
NVIDIA B200
71.7µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [152, 4096]
NVIDIA B200
71.8µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [200, 4096]
NVIDIA B200
71.8µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [168, 4096]
NVIDIA B200
71.8µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [176, 4096]
NVIDIA B200
71.8µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [88, 4096]
NVIDIA B200
71.8µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [192, 4096]
NVIDIA B200
71.9µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [208, 4096]
NVIDIA B200
71.9µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [184, 4096]
NVIDIA B200
72.0µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [96, 4096]
NVIDIA B200
72.4µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [216, 4096]
NVIDIA B200
72.4µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [104, 4096]
NVIDIA B200
72.5µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [112, 4096]
NVIDIA B200
72.8µs
#3 of 5
2025-10-16
GEMM n6144 k4096fp16 · [224, 4096]
NVIDIA B200
73.0µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [232, 4096]
NVIDIA B200
73.2µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [120, 4096]
NVIDIA B200
73.3µs
#4 of 5
2025-10-16
GEMM n6144 k4096fp16 · [240, 4096]
NVIDIA B200
73.6µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [128, 4096]
NVIDIA B200
73.7µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [248, 4096]
NVIDIA B200
73.9µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [256, 4096]
NVIDIA B200
74.1µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [972, 4096]
NVIDIA B200
77.8µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [2053, 4096]
NVIDIA B200
148.4µs
#5 of 6
2025-10-16
GEMM n6144 k4096fp16 · [2379, 4096]
NVIDIA B200
183.3µs
#4 of 6
2025-10-16
GEMM n6144 k4096fp16 · [8192, 4096]
NVIDIA B200
546.1µs
#5 of 6
2025-10-16

Reported · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:70f0256d7383fbab2d025fb6a4f788fd2e49461f400574ef2dd62168cc883f66
license declaredApache-2.0
license concludedApache-2.0
authorsclaude-opus-4-1-20250805
imported2026-08-20

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

mmaaccumulator = tl.dot(a, tl.trans(b), accumulator)
tile-k = 32BLOCK_K = 32
tile-m = 128BLOCK_M = 128
tile-n = 128BLOCK_N = 128

Kernel source

main.py115 lines
import torch
import triton
import triton.language as tl
import math

@triton.jit
def gemm_kernel(
    a_ptr, b_ptr, c_ptr,
    M, N, K,
    stride_am, stride_ak,
    stride_bn, stride_bk,
    stride_cm, stride_cn,
    BLOCK_M: tl.constexpr,
    BLOCK_N: tl.constexpr,
    BLOCK_K: tl.constexpr,
    GROUP_M: tl.constexpr,
):
    pid = tl.program_id(0)
    num_pid_m = tl.cdiv(M, BLOCK_M)
    num_pid_n = tl.cdiv(N, BLOCK_N)
    num_pid_in_group = GROUP_M * num_pid_n
    group_id = pid // num_pid_in_group
    first_pid_m = group_id * GROUP_M
    group_size_m = min(num_pid_m - first_pid_m, GROUP_M)
    pid_m = first_pid_m + ((pid % num_pid_in_group) % group_size_m)
    pid_n = (pid % num_pid_in_group) // group_size_m

    offs_am = pid_m * BLOCK_M + tl.arange(0, BLOCK_M)
    offs_bn = pid_n * BLOCK_N + tl.arange(0, BLOCK_N)
    offs_k = tl.arange(0, BLOCK_K)
    
    a_ptrs = a_ptr + (offs_am[:, None] * stride_am + offs_k[None, :] * stride_ak)
    b_ptrs = b_ptr + (offs_bn[:, None] * stride_bn + offs_k[None, :] * stride_bk)

    accumulator = tl.zeros((BLOCK_M, BLOCK_N), dtype=tl.float32)
    
    for k in range(0, tl.cdiv(K, BLOCK_K)):
        mask_k = (k * BLOCK_K + offs_k) < K
        
        a = tl.load(a_ptrs, mask=(offs_am[:, None] < M) & mask_k[None, :], other=0.0)
        b = tl.load(b_ptrs, mask=(offs_bn[:, None] < N) & mask_k[None, :], other=0.0)
        
        accumulator = tl.dot(a, tl.trans(b), accumulator)
        
        a_ptrs += BLOCK_K * stride_ak
        b_ptrs += BLOCK_K * stride_bk

    offs_cm = pid_m * BLOCK_M + tl.arange(0, BLOCK_M)
    offs_cn = pid_n * BLOCK_N + tl.arange(0, BLOCK_N)
    c_ptrs = c_ptr + stride_cm * offs_cm[:, None] + stride_cn * offs_cn[None, :]
    c_mask = (offs_cm[:, None] < M) & (offs_cn[None, :] < N)
    
    c = accumulator.to(tl.float16)
    tl.store(c_ptrs, c, mask=c_mask)

def run(A, B):
    # Handle device management
    original_device_A = A.device
    original_device_B = B.device
    
    if not A.is_cuda:
        if torch.cuda.is_available():
            A = A.cuda()
        else:
            raise RuntimeError("CUDA is not available for GPU tensor operations")
    
    if not B.is_cuda:
        if torch.cuda.is_available():
            B = B.cuda()
        else:
            raise RuntimeError("CUDA is not available for GPU tensor operations")
    
    # Get dimensions
    M, K_A = A.shape
    N, K_B = B.shape
    
    assert K_A == K_B, f"Dimension mismatch: A has K={K_A}, B has K={K_B}"
    K = K_A
    
    # Ensure inputs are float16
    A = A.to(torch.float16)
    B = B.to(torch.float16)
    
    # Allocate output
    C = torch.empty((M, N), device=A.device, dtype=torch.float16)
    
    # Configure kernel parameters for B200
    BLOCK_M = 128
    BLOCK_N = 128
    BLOCK_K = 32
    GROUP_M = 8
    
    # Calculate grid
    num_blocks = triton.cdiv(M, BLOCK_M) * triton.cdiv(N, BLOCK_N)
    
    # Launch kernel
    gemm_kernel[(num_blocks,)](
        A, B, C,
        M, N, K,
        A.stride(0), A.stride(1),
        B.stride(0), B.stride(1),
        C.stride(0), C.stride(1),
        BLOCK_M=BLOCK_M,
        BLOCK_N=BLOCK_N,
        BLOCK_K=BLOCK_K,
        GROUP_M=GROUP_M,
    )
    
    # Move result back to original device
    if original_device_A.type == 'cpu' and original_device_B.type == 'cpu':
        C = C.cpu()
    elif original_device_A != C.device:
        C = C.to(original_device_A)
    
    return C
scrolls · 115 lines total

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reported

JSON