gemini-2.5-pro / cudad4c20e
gemini-2.5-pro_cuda_d4c20e · gemini-2.5-pro · cuda · Apache-2.0
Use it
Vendorable · source mirrored · Apache-2.0View source →
No package. Vendor the mirrored source: 58 lines, Apache-2.0, pinned at da91508.
main.cpp
curl "https://kernelindex.com/api/v1/implementations/flashinfer-gemini-2-5-pro-cuda-d4c20e?include=source"interfacecuda
revisionda915083d4c7
symbolrun
pathmain.cpp
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16
Benchmark evidence
43 measurements across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Show all 43 measurements ›Showing all 43 measurements ⌄
Reproduction-ready · How evidence levels are derived →
Source and license
sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:a16718cb83db3fae24eb266924c7ee9422e77434ffaf5929e7a3cb06aa7dfa83
license declaredApache-2.0
license concludedApache-2.0
authorsgemini-2.5-pro
imported2026-08-20
Kernel source
main.cpp58 lines
#include "kernel.h"
#include <torch/extension.h>
#include <vector>
// Constants defined by the GEMM specification
constexpr int64_t N_DIM = 28672;
constexpr int64_t K_DIM = 4096;
/**
* @brief Python-bindable entry point for the GEMM operation.
*
* This function acts as a C++ interface between Python (PyTorch) and the CUDA
* kernel launcher. It performs extensive input validation, allocates the output
* tensor, and calls the CUDA implementation.
*
* @param A A PyTorch tensor representing matrix A with shape [M, 4096] and dtype float16.
* @param B A PyTorch tensor representing matrix B with shape [28672, 4096] and dtype float16.
* @return A new PyTorch tensor C, the result of A @ B.T, with shape [M, 28672] and dtype float16.
*/
torch::Tensor run(torch::Tensor A, torch::Tensor B) {
// --- Input Validation ---
TORCH_CHECK(A.dim() == 2, "Input tensor A must be 2-dimensional");
TORCH_CHECK(B.dim() == 2, "Input tensor B must be 2-dimensional");
TORCH_CHECK(A.is_cuda() && B.is_cuda(), "Input tensors must be on the same CUDA device");
TORCH_CHECK(A.device() == B.device(), "Input tensors must be on the same CUDA device");
TORCH_CHECK(A.scalar_type() == torch::kFloat16, "Input tensor A must have dtype float16");
TORCH_CHECK(B.scalar_type() == torch::kFloat16, "Input tensor B must have dtype float16");
TORCH_CHECK(A.size(1) == K_DIM, "Input tensor A must have K=", K_DIM, ", but got ", A.size(1));
TORCH_CHECK(B.size(0) == N_DIM, "Input tensor B must have N=", N_DIM, ", but got ", B.size(0));
TORCH_CHECK(B.size(1) == K_DIM, "Input tensor B must have K=", K_DIM, ", but got ", B.size(1));
TORCH_CHECK(A.size(1) == B.size(1), "Inner dimensions of A and B must match (K dimension)");
TORCH_CHECK(A.is_contiguous(), "Input tensor A must be contiguous");
TORCH_CHECK(B.is_contiguous(), "Input tensor B must be contiguous");
// --- Tensor Allocation ---
const int64_t M = A.size(0);
const auto C_shape = std::vector<int64_t>{M, N_DIM};
// Create the output tensor C on the same device and with the same dtype as the inputs.
torch::Tensor C = torch::empty(C_shape, A.options());
// --- Kernel Execution ---
// Launch the CUDA kernel through the host wrapper function.
gemm_n28672_k4096_launch(A, B, C);
return C;
}
// --- Pybind11 Module Definition ---
// This macro creates the Python module and binds the C++ 'run' function
// so it can be called from Python.
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
m.def("run", &run, "gemm_n28672_k4096(A, B) CUDA implementation. Computes C = A @ B.T.");
}scrolls · 58 lines total
Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0
Best evidence level for this revision: reproducible
JSON