Skip to content
KernelIndex
Search⌘K

gemini-2.5-pro / cudad4c20e

gemini-2.5-pro_cuda_d4c20e · gemini-2.5-pro · cuda · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 58 lines, Apache-2.0, pinned at da91508.

main.cpp
curl "https://kernelindex.com/api/v1/implementations/flashinfer-gemini-2-5-pro-cuda-d4c20e?include=source"
interfacecuda
revisionda915083d4c7
symbolrun
pathmain.cpp
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

43 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
GEMM n28672 k4096fp16 · [16, 4096]
NVIDIA B200
409.5µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [35, 4096]
NVIDIA B200
413.8µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [15, 4096]
NVIDIA B200
414.0µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [64, 4096]
NVIDIA B200
414.1µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [72, 4096]
NVIDIA B200
415.2µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [56, 4096]
NVIDIA B200
415.9µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [120, 4096]
NVIDIA B200
418.6µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [1, 4096]
NVIDIA B200
418.6µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [80, 4096]
NVIDIA B200
418.6µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [70, 4096]
NVIDIA B200
418.9µs
#8 of 8
2025-10-16
Show all 43 measurements ›
GEMM n28672 k4096fp16 · [144, 4096]
NVIDIA B200
419.3µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [200, 4096]
NVIDIA B200
420.7µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [96, 4096]
NVIDIA B200
422.4µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [176, 4096]
NVIDIA B200
422.9µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [88, 4096]
NVIDIA B200
425.7µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [152, 4096]
NVIDIA B200
426.8µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [248, 4096]
NVIDIA B200
428.7µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [8, 4096]
NVIDIA B200
444.3µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [128, 4096]
NVIDIA B200
444.6µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [4, 4096]
NVIDIA B200
444.7µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [232, 4096]
NVIDIA B200
446.4µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [7, 4096]
NVIDIA B200
446.7µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [2, 4096]
NVIDIA B200
447.0µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [40, 4096]
NVIDIA B200
448.0µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [240, 4096]
NVIDIA B200
457.3µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [256, 4096]
NVIDIA B200
482.0µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [136, 4096]
NVIDIA B200
507.4µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [48, 4096]
NVIDIA B200
511.5µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [184, 4096]
NVIDIA B200
524.3µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [2053, 4096]
NVIDIA B200
819.4µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [160, 4096]
NVIDIA B200
891.3µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [168, 4096]
NVIDIA B200
910.7µs
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [2379, 4096]
NVIDIA B200
930.6µs
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [32, 4096]
NVIDIA B200
1.06ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [224, 4096]
NVIDIA B200
1.07ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [216, 4096]
NVIDIA B200
1.07ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [192, 4096]
NVIDIA B200
1.08ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [104, 4096]
NVIDIA B200
1.11ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [208, 4096]
NVIDIA B200
1.16ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [24, 4096]
NVIDIA B200
1.17ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [112, 4096]
NVIDIA B200
1.22ms
#8 of 8
2025-10-16
GEMM n28672 k4096fp16 · [972, 4096]
NVIDIA B200
1.33ms
#7 of 8
2025-10-16
GEMM n28672 k4096fp16 · [8192, 4096]
NVIDIA B200
1.66ms
#4 of 8
2025-10-16

Reproduction-ready · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:a16718cb83db3fae24eb266924c7ee9422e77434ffaf5929e7a3cb06aa7dfa83
license declaredApache-2.0
license concludedApache-2.0
authorsgemini-2.5-pro
imported2026-08-20

Kernel source

main.cpp58 lines
#include "kernel.h"
#include <torch/extension.h>
#include <vector>

// Constants defined by the GEMM specification
constexpr int64_t N_DIM = 28672;
constexpr int64_t K_DIM = 4096;

/**
 * @brief Python-bindable entry point for the GEMM operation.
 *
 * This function acts as a C++ interface between Python (PyTorch) and the CUDA
 * kernel launcher. It performs extensive input validation, allocates the output
 * tensor, and calls the CUDA implementation.
 *
 * @param A A PyTorch tensor representing matrix A with shape [M, 4096] and dtype float16.
 * @param B A PyTorch tensor representing matrix B with shape [28672, 4096] and dtype float16.
 * @return A new PyTorch tensor C, the result of A @ B.T, with shape [M, 28672] and dtype float16.
 */
torch::Tensor run(torch::Tensor A, torch::Tensor B) {
    // --- Input Validation ---
    TORCH_CHECK(A.dim() == 2, "Input tensor A must be 2-dimensional");
    TORCH_CHECK(B.dim() == 2, "Input tensor B must be 2-dimensional");

    TORCH_CHECK(A.is_cuda() && B.is_cuda(), "Input tensors must be on the same CUDA device");
    TORCH_CHECK(A.device() == B.device(), "Input tensors must be on the same CUDA device");

    TORCH_CHECK(A.scalar_type() == torch::kFloat16, "Input tensor A must have dtype float16");
    TORCH_CHECK(B.scalar_type() == torch::kFloat16, "Input tensor B must have dtype float16");

    TORCH_CHECK(A.size(1) == K_DIM, "Input tensor A must have K=", K_DIM, ", but got ", A.size(1));
    TORCH_CHECK(B.size(0) == N_DIM, "Input tensor B must have N=", N_DIM, ", but got ", B.size(0));
    TORCH_CHECK(B.size(1) == K_DIM, "Input tensor B must have K=", K_DIM, ", but got ", B.size(1));
    TORCH_CHECK(A.size(1) == B.size(1), "Inner dimensions of A and B must match (K dimension)");

    TORCH_CHECK(A.is_contiguous(), "Input tensor A must be contiguous");
    TORCH_CHECK(B.is_contiguous(), "Input tensor B must be contiguous");

    // --- Tensor Allocation ---
    const int64_t M = A.size(0);
    const auto C_shape = std::vector<int64_t>{M, N_DIM};
    
    // Create the output tensor C on the same device and with the same dtype as the inputs.
    torch::Tensor C = torch::empty(C_shape, A.options());

    // --- Kernel Execution ---
    // Launch the CUDA kernel through the host wrapper function.
    gemm_n28672_k4096_launch(A, B, C);

    return C;
}

// --- Pybind11 Module Definition ---
// This macro creates the Python module and binds the C++ 'run' function
// so it can be called from Python.
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
    m.def("run", &run, "gemm_n28672_k4096(A, B) CUDA implementation. Computes C = A @ B.T.");
}
scrolls · 58 lines total

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reproducible

JSON