Recursive
Fused attention QK matmul scale mask softmax backward · suite of 16 cases · mean latency · run 01a0352a-1d9e…
●passedReported evidence
Primary measurement
144.6µsmean
Rank 2 in its comparison group · source-native comparison · observed 2026-06-07
SOL 0.6941 · 1.93× reference · 16/16 cases faster
Identity
implementationRecursive
projectRecursive
revisionsubmission-6709
workloadsuite of 16 cases · mean latency
comparison keysha256:92bc0be5b5ce9f52…
sourceNVIDIA SOL-ExecBench
external idsubmission/6709
sha256:7f6df4bbbdbeb0960fc8acc…
Correctness
Marked passed by the source; the correctness policy was not published.
Workload
definition comparatorsol_execbench_eval
descriptionCorrectness must pass on every case in the SOL evaluation stack.
Measurements
latency · mean144.6 µs
Protocol
harnessSOL-ExecBench evaluation stack v1.0
timerharness_reported
primaryStatisticmean
correctnessReferencedefinition reference implementation
comparabilityFamilysol_execbench_leaderboard
Environment
gpuNVIDIA B200 (sm_100)
Artifacts
No artifacts published with this run.
Replications and notes
No attestations yet.
Community attestations. They never change the evidence level; only a KernelIndex-controlled rerun does.
Add a reproduction or note
Canonical manifest
Show manifestHide manifest
{
"run": {
"kind": "BenchmarkRun",
"spec": {
"status": "passed",
"timing": {
"latencyNs": {
"mean": 144571
},
"primaryStatistic": "mean"
},
"observedAt": "2026-06-07T12:03:35.566Z",
"sourceNative": {
"source": "sol-execbench",
"metrics": {
"sol_score": 0.694123,
"latency_ms": 0.144571,
"avg_speedup": 1.9338,
"fast_1_count": 16,
"fast_1_total": 16
},
"benchmark": "060_fused_attention_qk_matmul_scale_mask_softmax_backward",
"externalId": "6709"
},
"protocolDigest": "sha256:2964c5c4ed9cb0dcdc0f74674035bc71bf5e66620939e38441060f76434ed3e5",
"workloadDigest": "sha256:b34041279b11ad4d663ebcd90e679a540745fbdc6b7f9e5b5eb70dd3ae256f84",
"environmentDigest": "sha256:5296c820e166474e590abb1beaf7728cdfe0a4bde9bf55f8966a40f4f5401621",
"implementationDigest": "sha256:ecd84fcaa9898a6b66832e03173fd9302f5296467b75b3a3e989e4de25fe8378"
},
"metadata": {
"name": "sol-submission-6709",
"title": "Recursive · 060_fused_attention_qk_matmul_scale_mask_softmax_backward · B200",
"labels": {
"worker": "B200-ext-prod-2527"
}
},
"apiVersion": "kernelindex.dev/v1alpha1"
},
"protocol": {
"kind": "BenchmarkProtocol",
"spec": {
"harness": {
"name": "SOL-ExecBench evaluation stack",
"version": "v1.0"
},
"correctness": {
"reference": "definition reference implementation",
"comparator": "sol_execbench_eval"
},
"measurement": {
"timer": "harness_reported",
"primaryStatistic": "mean"
},
"comparability": {
"notes": "Suite-mean latency and SOL-Score as published by the SOL-ExecBench leaderboard. Comparable only within one definition, GPU type, and evaluation stack version.",
"family": "sol_execbench_leaderboard"
}
},
"metadata": {
"name": "sol-execbench-leaderboard-v1-0",
"title": "SOL-ExecBench evaluation stack v1.0"
},
"apiVersion": "kernelindex.dev/v1alpha1"
},
"environment": {
"kind": "ExecutionEnvironment",
"spec": {
"hardware": {
"vendor": "nvidia",
"product": "NVIDIA B200",
"architecture": "sm_100"
},
"software": {
"libraries": {
"evaluation_stack": "v1.0"
}
}
},
"metadata": {
"name": "sol-execbench-b200-v1-0",
"title": "SOL-ExecBench NVIDIA B200, stack v1.0"
},
"apiVersion": "kernelindex.dev/v1alpha1"
}
}Cite this record (permalink, digest, access date)
Report an issue with this run
Published 2026-08-24 · NVIDIA SOL-ExecBenchAll results for Fused attention QK matmul scale mask softmax backward →JSON