flashinfer / wrapper123ca6
flashinfer_wrapper_123ca6 · FlashInfer-Bench baselines · python · Apache-2.0
Use it
Vendorable · source mirrored · Apache-2.0View source →
No package. Vendor the mirrored source: 35 lines, Apache-2.0, pinned at da91508.
main.py
curl "https://kernelindex.com/api/v1/implementations/flashinfer-flashinfer-wrapper-123ca6?include=source"interfacepython
revisionda915083d4c7
symbolrun
pathmain.py
Compatibility
declared hardwareNVIDIA B200
architecturesunknown
dtypesbf16, fp32, int64
Benchmark evidence
No published measurement for this revision.
No evidence · How evidence levels are derived →
Source and license
sourcehttps://huggingface.co/datasets/flashinfer-ai/flashinfer-trace
commitda915083d4c7c5e61aa3005e3d17ae488e0fc71c
revision digestsha256:0317a1cfd72511ed4b7b5ffbf5245b28dacfa08c2103e37f1e2fac50d80b8281
license declaredApache-2.0
license concludedApache-2.0
authorsbaseline
imported2026-08-16
Kernel source
main.py35 lines
import torch
import torch.nn.functional as F
from .gdn_blackwell.gdn import chunk_gated_delta_rule
def run(q, k, v, state, A_log, a, dt_bias, b, cu_seqlens, scale):
x = a.float() + dt_bias.float()
# Gate values in log-space: log(exp(-exp(A_log) * softplus(x))) = -exp(A_log) * softplus(x)
g = -torch.exp(A_log.float()) * F.softplus(x)
beta = torch.sigmoid(b.float())
varlen = cu_seqlens is not None and q.dim() == 3
if varlen:
q = q.unsqueeze(0)
k = k.unsqueeze(0)
v = v.unsqueeze(0)
output, new_state = chunk_gated_delta_rule(
q=q,
k=k,
v=v,
g=g,
beta=beta,
scale=None,
initial_state=state,
output_final_state=True,
cu_seqlens=cu_seqlens,
use_qk_l2norm_in_kernel=False,
)
if varlen:
output = output.squeeze(0)
return output, new_state
scrolls · 35 lines total
Source code from the importing source · Apache-2.0
No published measurement for this revision
JSON