flashinfer / wrapperc3f8a1
flashinfer_wrapper_c3f8a1 · FlashInfer-Bench baselines · python · Apache-2.0
Use it
Vendorable · source mirrored · Apache-2.0View source →
No package. Vendor the mirrored source: 25 lines, Apache-2.0, pinned at da91508.
main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-flashinfer-wrapper-c3f8a1?include=source"interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
declared hardwareNVIDIA H100, NVIDIA H20, NVIDIA H200
architecturesunknown
dtypesbf16, fp32, int64
Benchmark evidence
No published measurement for this revision.
No evidence · How evidence levels are derived →
Source and license
sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:a042ece35b3e6830eadd5ba18958cb7cca01aa9448698971d02ae41607914035
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16
Kernel source
main.py25 lines
import torch
import torch.nn.functional as F
from flashinfer.gdn_prefill import chunk_gated_delta_rule
def run(q, k, v, state, A_log, a, dt_bias, b, cu_seqlens, scale):
x = a.float() + dt_bias.float()
g = torch.exp(-torch.exp(A_log.float()) * F.softplus(x))
beta = torch.sigmoid(b.float())
output, new_state = chunk_gated_delta_rule(
q=q,
k=k,
v=v,
g=g,
beta=beta,
scale=None,
initial_state=state,
output_final_state=True,
cu_seqlens=cu_seqlens,
use_qk_l2norm_in_kernel=False,
)
return output, new_state
Source code from the importing source · Apache-2.0
No published measurement for this revision
JSON