Skip to content
KernelIndex
Search⌘K

flashinfer / wrapper03f7b0

flashinfer_wrapper_03f7b0 · FlashInfer-Bench baselines · python · Apache-2.0

Use it

Vendorable · source mirrored · Apache-2.0View source →

No package. Vendor the mirrored source: 97 lines, Apache-2.0, pinned at da91508.

main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-flashinfer-wrapper-03f7b0?include=source"
interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA A100, NVIDIA B200, NVIDIA GeForce RTX 4090, NVIDIA H100, NVIDIA H20, NVIDIA H200
architecturesunknown
dtypesbf16, fp32, int32

Benchmark evidence

47 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=8
NVIDIA B200
15.1µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=208
NVIDIA B200
18.3µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=457
NVIDIA B200
18.4µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=108
NVIDIA B200
18.5µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=308
NVIDIA B200
18.6µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=408
NVIDIA B200
19.6µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=1857
NVIDIA B200
19.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=508
NVIDIA B200
20.5µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=608
NVIDIA B200
20.5µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=1208
NVIDIA B200
20.7µs
#1 of 4
2025-10-21
Show all 47 measurements ›
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=708
NVIDIA B200
22.4µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=808
NVIDIA B200
23.9µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=2757
NVIDIA B200
24.6µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=1008
NVIDIA B200
24.6µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=1908
NVIDIA B200
24.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=3557
NVIDIA B200
24.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=4357
NVIDIA B200
25.2µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=1108
NVIDIA B200
26.3µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=2408
NVIDIA B200
26.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [1, 16, 64] · num_kv_indices=2708
NVIDIA B200
26.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=7257
NVIDIA B200
30.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=8057
NVIDIA B200
30.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=5057
NVIDIA B200
30.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=6657
NVIDIA B200
30.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=5857
NVIDIA B200
30.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=8857
NVIDIA B200
30.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=9657
NVIDIA B200
36.0µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=12857
NVIDIA B200
36.3µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=10857
NVIDIA B200
36.6µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=9945
NVIDIA B200
36.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=16345
NVIDIA B200
36.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=14857
NVIDIA B200
36.9µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [16, 16, 64] · num_kv_indices=17257
NVIDIA B200
36.9µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=22745
NVIDIA B200
37.9µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=27545
NVIDIA B200
44.9µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=30745
NVIDIA B200
49.3µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=33945
NVIDIA B200
54.4µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=37145
NVIDIA B200
57.4µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=44845
NVIDIA B200
59.4µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=40345
NVIDIA B200
59.5µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=48045
NVIDIA B200
59.7µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=51245
NVIDIA B200
59.8µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=54445
NVIDIA B200
61.3µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=57645
NVIDIA B200
61.4µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=62345
NVIDIA B200
64.5µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=68745
NVIDIA B200
69.5µs
#1 of 4
2025-10-21
MLA paged decode h16 ckv512 kpe64 ps1bf16 · [64, 16, 64] · num_kv_indices=75145
NVIDIA B200
83.6µs
#1 of 4
2025-10-21

Reported · How evidence levels are derived →

Source and license

sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:d63f2eb0f0a06bf1c4e9dd1f35d4aebd4c281ea806f8598245b7382024e120b5
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16

Kernel source

main.py97 lines
import torch
import flashinfer

_WORKSPACE_SIZE_BYTES = 128 * 1024 * 1024
_workspace_cache = {}
_wrapper_cache = {}
_plan_state = {}


def _get_workspace(device):
    key = str(device)
    buffer = _workspace_cache.get(key)
    if buffer is None or buffer.device != device or buffer.numel() < _WORKSPACE_SIZE_BYTES:
        buffer = torch.empty(_WORKSPACE_SIZE_BYTES, dtype=torch.int8, device=device)
        _workspace_cache[key] = buffer
    return buffer


def _get_wrapper(key, device):
    wrapper = _wrapper_cache.get(key)
    if wrapper is None:
        workspace = _get_workspace(device)
        wrapper = flashinfer.mla.BatchMLAPagedAttentionWrapper(workspace)
        _wrapper_cache[key] = wrapper
    return wrapper


def run(q_nope, q_pe, ckv_cache, kpe_cache, kv_indptr, kv_indices, sm_scale):
    batch_size, num_qo_heads, head_dim_ckv = q_nope.shape
    head_dim_kpe = q_pe.shape[-1]
    page_size = ckv_cache.shape[1]
    len_indptr = kv_indptr.shape[0]
    num_kv_indices = kv_indices.shape[0]

    device = q_nope.device
    wrapper_key = (
        str(device),
        num_qo_heads,
        head_dim_ckv,
        head_dim_kpe,
        page_size,
        q_nope.dtype,
        q_pe.dtype,
        ckv_cache.dtype,
        kpe_cache.dtype,
    )

    wrapper = _get_wrapper(wrapper_key, device)
    state = _plan_state.get(wrapper_key)

    needs_plan = True
    if state is not None:
        needs_plan = (
            state.get("batch_size") != batch_size
            or state.get("len_indptr") != len_indptr
            or state.get("num_kv_indices") != num_kv_indices
            or state.get("sm_scale") != sm_scale
            or state.get("kv_indptr_ptr") != kv_indptr.data_ptr()
            or state.get("kv_indices_ptr") != kv_indices.data_ptr()
        )

    if needs_plan:
        qo_indptr = torch.arange(0, batch_size + 1, dtype=torch.int32, device=device)
        kv_len_arr = (kv_indptr[1:] - kv_indptr[:-1]).to(torch.int32)
        wrapper.plan(
            qo_indptr=qo_indptr,
            kv_indptr=kv_indptr,
            kv_indices=kv_indices,
            kv_len_arr=kv_len_arr,
            num_heads=num_qo_heads,
            head_dim_ckv=head_dim_ckv,
            head_dim_kpe=head_dim_kpe,
            page_size=page_size,
            causal=False,
            sm_scale=sm_scale,
            q_data_type=q_nope.dtype,
            kv_data_type=ckv_cache.dtype,
        )
        _plan_state[wrapper_key] = {
            "batch_size": batch_size,
            "len_indptr": len_indptr,
            "num_kv_indices": num_kv_indices,
            "sm_scale": sm_scale,
            "kv_indptr_ptr": kv_indptr.data_ptr(),
            "kv_indices_ptr": kv_indices.data_ptr(),
        }

    output, lse = wrapper.run(
        q_nope,
        q_pe,
        ckv_cache,
        kpe_cache,
        return_lse=True,
    )

    return output, lse
scrolls · 97 lines total

Source code from FlashInfer-Bench (flashinfer-ai/flashinfer-trace) · Apache-2.0

Best evidence level for this revision: reported

JSON