Skip to content
KernelIndex
Search⌘K

claude-opus-4-1 / cuda1970e7

claude-opus-4-1_cuda_1970e7 · claude-opus-4-1-20250805 · cuda · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 63 lines, Apache-2.0, pinned at da91508.

main.cpp
curl "https://kernelindex.com/api/v1/implementations/flashinfer-claude-opus-4-1-cuda-1970e7?include=source"
interfacecuda
revisionda915083d4c7
symbolrun
pathmain.cpp
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

22 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n4096 k4096fp16 · [24, 4096]
NVIDIA B200
279.7µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [32, 4096]
NVIDIA B200
279.7µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [15, 4096]
NVIDIA B200
279.9µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [8, 4096]
NVIDIA B200
280.0µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [16, 4096]
NVIDIA B200
280.0µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [7, 4096]
NVIDIA B200
280.0µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [1, 4096]
NVIDIA B200
280.1µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [2, 4096]
NVIDIA B200
280.1µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [4, 4096]
NVIDIA B200
280.1µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [56, 4096]
NVIDIA B200
475.6µs
#7 of 8
2025-10-16
Show all 22 measurements ›
GEMM n4096 k4096fp16 · [48, 4096]
NVIDIA B200
475.7µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [35, 4096]
NVIDIA B200
475.8µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [40, 4096]
NVIDIA B200
476.2µs
#7 of 8
2025-10-16
GEMM n4096 k4096fp16 · [80, 4096]
NVIDIA B200
817.8µs
#8 of 8
2025-10-16
GEMM n4096 k4096fp16 · [208, 4096]
NVIDIA B200
821.0µs
#9 of 9
2025-10-16
GEMM n4096 k4096fp16 · [224, 4096]
NVIDIA B200
846.5µs
#8 of 8
2025-10-16
GEMM n4096 k4096fp16 · [96, 4096]
NVIDIA B200
846.9µs
#8 of 8
2025-10-16
GEMM n4096 k4096fp16 · [240, 4096]
NVIDIA B200
869.4µs
#8 of 8
2025-10-16
GEMM n4096 k4096fp16 · [112, 4096]
NVIDIA B200
880.0µs
#8 of 8
2025-10-16
GEMM n4096 k4096fp16 · [256, 4096]
NVIDIA B200
887.5µs
#9 of 9
2025-10-16
GEMM n4096 k4096fp16 · [128, 4096]
NVIDIA B200
911.6µs
#9 of 9
2025-10-16
GEMM n4096 k4096fp16 · [8192, 4096]
NVIDIA B200
5.40ms
#8 of 8
2025-10-16

Reproduction-ready · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:97b87e562d6fecf308c954cc6602a86a0672190908046ae9b46f8e4444966d8a
license declaredApache-2.0
license concludedApache-2.0
authorsclaude-opus-4-1-20250805
imported2026-08-20

Kernel source

main.cpp63 lines
#include <torch/extension.h>
#include <cuda_runtime.h>
#include <cuda_fp16.h>
#include "kernel.h"
#include <ATen/cuda/CUDAContext.h>
#include <c10/cuda/CUDAGuard.h>

torch::Tensor run(torch::Tensor A, torch::Tensor B) {
    // Input validation
    TORCH_CHECK(A.dtype() == torch::kFloat16, "A must be float16");
    TORCH_CHECK(B.dtype() == torch::kFloat16, "B must be float16");
    TORCH_CHECK(A.is_cuda(), "A must be a CUDA tensor");
    TORCH_CHECK(B.is_cuda(), "B must be a CUDA tensor");
    TORCH_CHECK(A.is_contiguous(), "A must be contiguous");
    TORCH_CHECK(B.is_contiguous(), "B must be contiguous");
    
    // Dimension validation
    TORCH_CHECK(A.dim() == 2, "A must be 2D");
    TORCH_CHECK(B.dim() == 2, "B must be 2D");
    
    const int64_t M = A.size(0);
    const int64_t K_A = A.size(1);
    const int64_t N = B.size(0);
    const int64_t K_B = B.size(1);
    
    TORCH_CHECK(K_A == 4096, "A's K dimension must be 4096");
    TORCH_CHECK(N == 4096, "B's N dimension must be 4096");
    TORCH_CHECK(K_B == 4096, "B's K dimension must be 4096");
    
    // Set the CUDA device
    c10::cuda::CUDAGuard device_guard(A.device());
    
    // Create output tensor
    auto options = torch::TensorOptions()
        .dtype(torch::kFloat16)
        .device(A.device())
        .requires_grad(false);
    torch::Tensor C = torch::empty({M, N}, options);
    
    // Get CUDA stream
    cudaStream_t stream = at::cuda::getCurrentCUDAStream();
    
    // Get raw pointers
    const half* A_ptr = reinterpret_cast<const half*>(A.data_ptr<at::Half>());
    const half* B_ptr = reinterpret_cast<const half*>(B.data_ptr<at::Half>());
    half* C_ptr = reinterpret_cast<half*>(C.data_ptr<at::Half>());
    
    // Launch kernel
    launch_gemm_kernel(A_ptr, B_ptr, C_ptr, static_cast<int>(M), stream);
    
    // Check for errors
    cudaError_t err = cudaGetLastError();
    if (err != cudaSuccess) {
        TORCH_CHECK(false, "CUDA kernel launch failed: ", cudaGetErrorString(err));
    }
    
    return C;
}

PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
    m.def("run", &run, "Optimized GEMM for N=4096, K=4096",
          py::arg("A"), py::arg("B"));
}
scrolls · 63 lines total

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reproducible

JSON