submission 29936
AMD MLA decode · dq = 1536 · dim = 7168 · seed = 5291 · prefill = 6144 · batchsize = 128 · bf16 · run 01a006cd-a6ea…
●passedReported evidence
Not re-observed recently.
Primary measurement
2.02msmean of 3
Rank 18 in its comparison group · source-native comparison · observed 2025-06-01
mean ns 2017881.6666666667 · leaderboard score s 0.0020178816666666668
Identity
implementationsubmission 29936
projectKernelBot user 44
revisionsubmission-29936
workloaddq = 1536 · dim = 7168 · seed = 5291 · prefill = 6144 · batchsize = 128 · bf16
comparison keysha256:06ea84d769fae242…
sourceGPU MODE KernelBot
external idsubmission/29936/benchmark/0
sha256:564e59aeda01f9d2a190a72…
Correctness
Marked passed by the source; the correctness policy was not published.
Workload
dq1536
dim7168
seed5291
prefill6144
batchsize128
inputbf16 [128, 1, 7168]
kv_cachebf16 [128, 6144, 576]
definition comparatorkernelbot_eval
Measurements
latency · mean2.02 ms · n=3
latency · std9.24 µs · n=3
Protocol
harnessreference-kernels eval.py
timerwall_clock
primaryStatisticmean
correctnessReferenceproblem reference implementation
comparabilityFamilygpumode_kernelbot
Environment
gpuAMD Instinct MI300X (gfx942)
Artifacts
No artifacts published with this run.
Replications and notes
No attestations yet.
Community attestations. They never change the evidence level; only a KernelIndex-controlled rerun does.
Add a reproduction or note
Canonical manifest
Show manifestHide manifest
{
"run": {
"kind": "BenchmarkRun",
"spec": {
"status": "passed",
"timing": {
"samples": 3,
"latencyNs": {
"mean": 2017882
},
"primaryStatistic": "mean"
},
"observedAt": "2025-06-01T12:05:22.137Z",
"measurements": [
{
"unit": "ns",
"value": 9239.523490599142,
"metric": "latency",
"statistic": "std",
"sampleCount": 3
}
],
"sourceNative": {
"source": "gpumode-kernelbot",
"metrics": {
"mean_ns": 2017881.6666666667,
"leaderboard_score_s": 0.0020178816666666668
},
"benchmark": "amd-mla-decode/0",
"externalId": "submission/29936"
},
"protocolDigest": "sha256:7c66da4b5866502eb10444a7ba013d0f6bdf1620c55847945008cad2fbf98aec",
"workloadDigest": "sha256:e7946ffcd6ffe752e8cfc217f571a6ca854a6e541b06627c9e719940b2ae9fc7",
"environmentDigest": "sha256:f722892449ed976f3e21b12df8b92b34573fc232883fec6a2f3da399a00abdd5",
"implementationDigest": "sha256:fd1d54ad258c889155bab680698d2d54a33be18763772596aec595854ed7644a"
},
"metadata": {
"name": "kernelbot-29936-batchsize128-dim7168-dq1536-prefill6144-seed5291",
"title": "amd-mla-decode · submission 29936 · case 0",
"labels": {
"torch": "2.8.0.dev20250530+rocm6.3",
"platform": "Linux-5.15.0-136-generic-x86_64-with-glibc2.35"
}
},
"apiVersion": "kernelindex.dev/v1alpha1"
},
"protocol": {
"kind": "BenchmarkProtocol",
"spec": {
"harness": {
"name": "reference-kernels eval.py",
"repository": "https://github.com/gpu-mode/reference-kernels"
},
"correctness": {
"reference": "problem reference implementation",
"comparator": "kernelbot_eval"
},
"measurement": {
"timer": "wall_clock",
"primaryStatistic": "mean"
},
"comparability": {
"notes": "Per-case wall-clock statistics as published in the KernelBot dataset. Comparable only within one leaderboard problem and GPU type; runner software versions ride each run's labels.",
"family": "gpumode_kernelbot"
}
},
"metadata": {
"name": "gpumode-kernelbot-eval",
"title": "GPU MODE KernelBot evaluation harness"
},
"apiVersion": "kernelindex.dev/v1alpha1"
},
"environment": {
"kind": "ExecutionEnvironment",
"spec": {
"hardware": {
"vendor": "amd",
"product": "AMD Instinct MI300X",
"architecture": "gfx942"
},
"software": {}
},
"metadata": {
"name": "gpumode-kernelbot-amd-instinct-mi300x",
"title": "KernelBot runner · AMD Instinct MI300X"
},
"apiVersion": "kernelindex.dev/v1alpha1"
}
}Cite this record (permalink, digest, access date)
Report an issue with this run
Published 2026-08-15 · GPU MODE KernelBot · June 9 Researcher Reciprocity License v1.0All results for AMD MLA decode →JSON