doubleAI
Competition author · sol-execbench/user/doubleAI · license unknown
Records held
81as of 2026-10-01 09:05 UTC
Entries
712
Runs
712
Is this you? Claim this project ›
A claim grants attribution and metadata maintenance, never the right to edit evidence. Evidence claims are reviewed by a maintainer.
Current records
Operation · workload
History
Record
Implementation
Hardware
Since
GQA ragged prefill causal h32 kv4 d128suite of 15 cases
7.94µs
0.9% faster · was 8.01 µs
B200
2026-07-02
Show all 40 records ›Showing all 40 records ⌄
NVFP4 vision temporal patch merge with projectionsuite of 16 cases
42.2µs
27.1% faster · was 57.9 µs
B200
2026-06-27
MLA paged prefill causal h16 ckv512 kpe64 ps1suite of 38 cases
38.4µs
16.4% faster · was 46.0 µs
B200
2026-06-29
Chunk gated delta rule linear attentionsuite of 16 cases
70.4µs
7.8% faster · was 76.3 µs
B200
2026-06-26
MoE layer complete forward with residualsuite of 16 cases
1.11ms
27.7% faster · was 1.54 ms
B200
2026-06-27
Fused final layer upsample with adaptive normsuite of 16 cases
3.49ms
14.0% faster · was 4.05 ms
B200
2026-05-21
SAM hq mask decoder IoU hypernetwork fusion backwardsuite of 16 cases
57.8µs
38.2% faster · was 93.6 µs
B200
2026-06-26
Seqlen-finetuned-reconstructed Hyena gating and output projectionsuite of 16 cases
61.2µs
7.2% faster · was 66.0 µs
B200
2026-06-26
Seqlen-finetuned-reconstructed Hyena complete forward blocksuite of 16 cases
105.7µs
21.8% faster · was 135.2 µs
B200
2026-06-26
Audio encoder to language model multimodal fusionsuite of 16 cases
85.8µs
32.3% faster · was 126.8 µs
B200
2026-06-26
Audio sinusoidal position embedding with conv projectionsuite of 16 cases
405.2µs
46.3% faster · was 754.8 µs
B200
2026-06-26
Audio encoder varlen attention with chunking backwardsuite of 16 cases
113.4µs
1.0% faster · was 114.5 µs
B200
2026-06-26
MoE expert batched execution with capacity factorsuite of 16 cases
2.46ms
12.3% faster · was 2.81 ms
B200
2026-06-26
MoE expert computation with weighted accumulationsuite of 16 cases
403.7µs
0.2% faster · was 404.6 µs
B200
2026-06-26
Multimodal RoPE position computation with grid based indexingsuite of 16 cases
4.00µs
31.7% faster · was 5.86 µs
B200
2026-06-25
Mamba2 fused intra chunk diagonal computationsuite of 14 cases
11.9µs
30.4% faster · was 17.1 µs
B200
2026-06-29
Vision 3d rotary embedding with spatial merge indexing backwardsuite of 16 cases
11.4µs
2.3% faster · was 11.6 µs
B200
2026-06-25
41 more records · Open in the records ledger →
Measured entries
Implementation
Operation
Runtime
Best latency
Evidence
Availability
Show all 40 entries ›Showing all 40 entries ⌄
unknown
4.00µs
Reported
license unknown · no source
unknown
4.05µs
Reported
license unknown · no source
Showing 40 of 712; the full list is in the API dossier below.
Record activity
Nov
Dec
Jan
Feb
Mar
107Apr
71May
114Jun
30Jul
Aug
Sep
Oct
Sources: NVIDIA SOL-ExecBench (2026-07-14)last observed 2026-07-14JSON