Hyra
Attention QK matmul with GQA repeat and scaling · suite of 16 cases · mean latency · run 01a0352b-a6ea…
●passedReported evidence
Primary measurement
12.8µsmean
Rank 5 in its comparison group · source-native comparison · observed 2026-07-16
SOL 0.5488 · 1.49× reference · 7/16 cases faster
Identity
implementationHyra
projectHyra
revisionsubmission-24583
workloadsuite of 16 cases · mean latency
comparison keysha256:786b450aaf9fa4cc…
sourceNVIDIA SOL-ExecBench
external idsubmission/24583
sha256:ce0d108d3c936be83d7ae49…
Correctness
Marked passed by the source; the correctness policy was not published.
Workload
definition comparatorsol_execbench_eval
descriptionCorrectness must pass on every case in the SOL evaluation stack.
Measurements
latency · mean12.8 µs
Protocol
harnessSOL-ExecBench evaluation stack v1.1
timerharness_reported
primaryStatisticmean
correctnessReferencedefinition reference implementation
comparabilityFamilysol_execbench_leaderboard
Environment
gpuNVIDIA B200 (sm_100)
Artifacts
No artifacts published with this run.
Replications and notes
No attestations yet.
Community attestations. They never change the evidence level; only a KernelIndex-controlled rerun does.
Add a reproduction or note
Canonical manifest
Show manifestHide manifest
{
"run": {
"kind": "BenchmarkRun",
"spec": {
"status": "passed",
"timing": {
"latencyNs": {
"mean": 12840
},
"primaryStatistic": "mean"
},
"observedAt": "2026-07-16T06:21:12.598Z",
"sourceNative": {
"source": "sol-execbench",
"metrics": {
"sol_score": 0.548826,
"latency_ms": 0.01284,
"avg_speedup": 1.4851,
"fast_1_count": 7,
"fast_1_total": 16
},
"benchmark": "049_attention_qk_matmul_with_gqa_repeat_and_scaling",
"externalId": "24583"
},
"protocolDigest": "sha256:6383f62490dad270ad2146a94f67b48fba5cfcbc0b9b180f02665c69cc4433a0",
"workloadDigest": "sha256:8285d2d943ed0f1540bef26f8ca4a2d98b09716769a522085a968bf6a188f4f1",
"environmentDigest": "sha256:f304afb2fedab4bbeadf8ac36a990ff517ab6e824dfe5a3ee0fe46259ff5e39f",
"implementationDigest": "sha256:bda5405d08a103b9ab61f1d962eea0bd5cdbf0bc95f06f6585f61a3241e0ba2e"
},
"metadata": {
"name": "sol-submission-24583",
"title": "Hyra · 049_attention_qk_matmul_with_gqa_repeat_and_scaling · B200",
"labels": {
"worker": "external-prod-b200-modal-1"
}
},
"apiVersion": "kernelindex.dev/v1alpha1"
},
"protocol": {
"kind": "BenchmarkProtocol",
"spec": {
"harness": {
"name": "SOL-ExecBench evaluation stack",
"version": "v1.1"
},
"correctness": {
"reference": "definition reference implementation",
"comparator": "sol_execbench_eval"
},
"measurement": {
"timer": "harness_reported",
"primaryStatistic": "mean"
},
"comparability": {
"notes": "Suite-mean latency and SOL-Score as published by the SOL-ExecBench leaderboard. Comparable only within one definition, GPU type, and evaluation stack version.",
"family": "sol_execbench_leaderboard"
}
},
"metadata": {
"name": "sol-execbench-leaderboard-v1-1",
"title": "SOL-ExecBench evaluation stack v1.1"
},
"apiVersion": "kernelindex.dev/v1alpha1"
},
"environment": {
"kind": "ExecutionEnvironment",
"spec": {
"hardware": {
"vendor": "nvidia",
"product": "NVIDIA B200",
"architecture": "sm_100"
},
"software": {
"libraries": {
"evaluation_stack": "v1.1"
}
}
},
"metadata": {
"name": "sol-execbench-b200-v1-1",
"title": "SOL-ExecBench NVIDIA B200, stack v1.1"
},
"apiVersion": "kernelindex.dev/v1alpha1"
}
}Cite this record (permalink, digest, access date)
Report an issue with this run
Published 2026-08-24 · NVIDIA SOL-ExecBenchAll results for Attention QK matmul with GQA repeat and scaling →JSON