Documentation
Get started
- Search rmsnorm B200 bf16. The result lists the fastest known implementations measured on that workload, in rank order.
- The same answer over REST:
curl "https://kernelindex.com/api/v1/search?q=rmsnorm%20B200%20bf16"
- Or in the terminal: ki search "rmsnorm B200 bf16" --json.
- Opening a result shows the code, the license, and the benchmark run behind the number. The install command and the mirrored source are available from the same page.
KernelIndex indexes GPU kernel benchmark results. Each number is imported from a named public source, retained as published, and compared only with runs that measured the same quantity under the same conditions. Workloads for which the index holds no adequate answer are listed under challenges.
Searching
Enter an operation name, then add hardware, dtype, or shape to narrow the request. When a query matches several operations, the most heavily measured one is used and the page states which; the remainder appear under All matches.
Query syntax
rmsnorm rmsnorm B200 bf16 [2048,4096] tokens=2048 model:deepseek-v3 op:004-gemm-n128-k2048 rmsnorm gpu:B200 dtype:bf16 shape:[2048,4096] framework=pytorch trust:verified
Filters take key:value or key=value. Keys: op family model gpu arch dtype shape layout framework language cuda trust license source installable tech, plus axis bindings such as tokens=2048. Workload and hardware filters determine what qualifies as an exact match; trust, license, and technique filters only remove rows from the listing. tech:tma keeps implementations whose mirrored source uses that technique (the traits are extracted by pattern and listed on each implementation page). A misspelled filter produces a correction hint. If the exact shape requested has not been measured, the answer presents the nearest measured cases on either side of it.
Views
Results are separated into four views, which are never combined: Exact (matches the request), Compatible (close, with the differences listed), Supported (support claimed, no measurement), and Other protocols (measured by a different method). Recommended orders by strength of evidence and Newest orders by date. Ranks retain their meaning under either ordering.
To cover an entire model at once, use the model view: select a model and a GPU to see the best known implementation for each operation, with any gaps stated. The same answer is served at /api/v1/models/{slug}?gpu=.
Reading a result
Each number is presented with three facts: what it may be compared with, how it ranks, and the level of evidence supporting it.
Comparability
Runs are compared only within a cohort, meaning runs that measured the same quantity by the same method: the same workload, protocol, environment, and correctness threshold. A rank is meaningful only within a single cohort. Two runs that share nothing but a GPU name or an operation name are not compared.
Ranking
Inside a cohort, runs are ordered by latency under the frozen ranking-v1 policy. Runs whose difference falls within measurement uncertainty share a rank, shown as N=. Each excluded run is given a reason code, such as RETRACTED or MISSING_PRIMARY_METRIC.
Headroom
Beside a cohort record the operation page states an estimated floor: the time the workload's declared tensors need to cross HBM once at the GPU's datasheet bandwidth, and, for GEMM and attention families, the time their arithmetic needs at the dense tensor-core peak. The larger of the two values is the floor, and the record's distance above it indicates how much room may remain. The figure is a coarse lower bound under headroom-v1. It is labeled basis: estimate wherever it appears and is not treated as evidence, since a kernel may sit well above the floor for legitimate reasons. No estimate is produced for a GPU absent from the datasheet table.
Evidence levels
Levels are derived from stored facts and cannot be chosen by a submitter. Whether a kernel can be used in practice, which depends on its license, install recipe, and hardware, is tracked separately; the fastest result and the fastest usable result are frequently different rows.
Run pages also collect community attestations: reproduced, could not reproduce, an environment note, or a regression observed, each with an optional measured value and evidence link. Attestations accumulate alongside the evidence. Only a rerun on a KernelIndex-controlled runner changes an evidence level.
Using a kernel
Every implementation page opens with a Use it section, which covers three cases.
- A package exists, so the install command can be copied directly. When the evidence records a measured release, the command pins to it, so the installed version is the one that was measured. A command that cannot be pinned is labeled unpinned and never counted as usable.
- No package exists, but the source is mirrored and can be vendored. Copy the file from the page, or run ki use <implementation> to write it locally with the commit, license, and digest recorded in a header comment.
- No public source exists, so the row provides benchmark evidence only and states as much.
Every run page provides a Cite this record action, giving a permalink, a digest, and an access date.
API and agents
The API returns the same answers as these pages. Reference: /docs/api; machine contract: /api/v1/openapi.json.
# search: a person in a browser, or:
curl "https://kernelindex.com/api/v1/search?q=rmsnorm%20B200%20bf16"
# structured resolution: an agent with an exact workload:
curl -X POST https://kernelindex.com/api/v1/resolve/kernel \
-H 'Content-Type: application/json' \
-d '{"operation":{"name":"rmsnorm","axes":{"tokens":2048}},
"environment":{"hardwareProduct":"B200","dtype":"bf16"}}'
# many workloads in one call (an agent planning every operation of a model):
curl -X POST https://kernelindex.com/api/v1/resolve/kernel/batch \
-H 'Content-Type: application/json' \
-d '{"requests":[
{"operation":{"name":"rmsnorm"},"environment":{"hardwareProduct":"B200"}},
{"operation":{"name":"gemm"},"environment":{"hardwareProduct":"B200"}}]}'
# evidence dossiers (same models as the pages):
curl https://kernelindex.com/api/v1/runs/<id-or-digest>
curl "https://kernelindex.com/api/v1/implementations/<slug>?include=source"
# records ledger, cursor-paginated:
curl "https://kernelindex.com/api/v1/records?limit=50"
# what the index learned since you last polled:
curl "https://kernelindex.com/api/v1/feed?since=2026-08-01T00:00:00Z"
# the ki CLI, from a checkout (not on npm yet):
git clone https://github.com/SamMausberg/KernelIndex && cd KernelIndex
pnpm install --frozen-lockfile
alias ki="node apps/cli/src/ki.ts"
ki search "gemm b200 nvfp4" --json | jq '.groups.exact[0]'
ki use <implementation> # vendor a mirrored kernel source locally
ki manifest digest my-run.yaml
# validate a submission or flat bench record and preview its placement:
curl -X POST https://kernelindex.com/api/v1/submissions/preview \
-H 'Content-Type: application/json' \
-d "{\"document\": $(jq -Rs . < record.json)}"
# bulk export (versioned, immutable, zstd JSONL):
curl -L https://kernelindex.com/api/v1/exports/catalog.jsonl.zst
# README badge: current records held by an implementation (SVG):
API keys. Reads require no key. A key from your account raises the daily quota. Send it as Authorization: Bearer ki_… (CLI: --api-key or $KI_API_KEY). Requests over quota return 429 with Retry-After. Keys are stored as hashes and can be revoked at any time.
Precedents
These answer two different questions. Resolve identifies which indexed implementation can serve a given workload as it stands. Precedents identifies which code to study before writing a new implementation: for a problem the index may not have seen, it returns the implementations most likely to carry transferable optimization ideas, ranked by transferability (same computation, same or adjacent GPU architecture, adjacent shape, record standing, shared techniques) with the reasons stated. It expresses a study priority and is not a benchmark ranking.
ki precedents --op gqa-paged-decode --gpu B200 --dtype bf16 tokens=4096
curl -X POST https://kernelindex.com/api/v1/precedents \
-H "Content-Type: application/json" \
-d '{"operation": {"family": "gqa-paged-attention"}, "environment": {"hardwareProduct": "B200"}}'Agents
Point an agent at /llms.txt, a single page listing the claims this index supports and every machine-readable surface. MCP is hosted, so one URL completes the setup:
{ "mcpServers": { "kernelindex": { "url": "https://kernelindex.com/mcp" } } }
claude mcp add --transport http kernelindex https://kernelindex.com/mcpOver stdio instead, from a checkout: node apps/mcp/src/server.ts (KI_API overrides the API base, KI_API_KEY raises the quota). The published package @kernelindex/mcp is not on npm yet, so prefer the hosted URL above. Both paths run the same eighteen read-only tools over the public REST API.
REST, the CLI, MCP, the bulk export, the Atom feed, and the change feed (GET /feed?since=) return the same answers as these pages.
Serving covers whole LLM deployments, a separate corpus with its own comparisons. Pick an objective (say, maximize tokens/s under p99 ttft_ms ≤ 450) and the cohort ranks under it; pick none and you get the trade-off frontier. API: POST /api/v1/resolve/serving; CLI: ki resolve serving --manifest req.yaml.
Contributing
Records
A record is the fastest eligible run within a single cohort. The ledger presents the append-only run history: a record may be beaten or retracted, but never edited. Any two runs can be placed side by side on compare. A source's own reference implementation that has not yet been challenged is treated as coverage rather than a record; the ledger hides these by default and labels them baseline · unbeaten.
Every run page provides a Report an issue action, which requires no account. An accepted report retracts or supersedes the record, and the history remains visible. The feed lists the changes recorded by the index over the preceding 30 days: record breaks, imports, corrections, and accepted claims. When signed in, Following narrows the feed to the cohorts, operations, projects, GPUs, and models you follow.
Every project has a page with the records it holds and every kernel it measured. Authors may claim their own: a GitHub-hosted project is claimed in a single step by the account that owns the repository path, and any other case proceeds through reviewed evidence. A claim confers attribution and does not confer any right to edit evidence.
To contribute evidence, validate a submission and preview its placement with ki submit record.yaml, then use --send with an API key. Contribute →
How we count
Four counters appear across the site and are not interchangeable. Each surface states which counter it reports, so that any two can be reconciled arithmetically.
- Records. One per comparison cohort: the fastest eligible run in it. A cohort is narrower than an operation, so records outnumber operations.
- Operations with ranked runs. Operations holding at least one eligible run. The homepage and the ledger state this one.
- Operations indexed. Every definition in the catalog, whether measured or not. The remainder is reported on browse as indexed without runs, which records a gap in coverage and implies nothing about performance.
- Browse rows. What browse lists. Fewer than the operations behind them: definitions reviewed as equivalent fold into one row, which states how many it absorbed. Their own pages and cohorts stay separate.
Record counts also carry the ledger snapshot they came from. Pages cache independently, so two surfaces can state counts a few minutes of imports apart; the dates say so rather than leaving it to be read as a contradiction.
Sources and licensing
Every result is imported from a named public source and shown as published. Each source's license and required credit is on Legal. Rights holders can have anything removed; contested records come down first, questions after. "Ranked" counts the runs every ranked surface counts; "indexed" is the raw published corpus, failed and superseded runs included.
Serving results are held separately from kernel results. The Configs column counts distinct launch configurations.
Known limitations
- No result in the index has been reproduced by KernelIndex. Each measurement is presented as its source published it, and no record has therefore reached the Verified evidence level.
- SOL-ExecBench leaderboard rows report an average across a kernel's full case suite rather than individual cases, and are not used to answer a request for one specific case.
- FlashInfer-Bench imports consist of library baselines taken at a pinned dataset revision, together with model-generated solutions. Each model-generated record is labeled as such and names the model that produced it.
- Liger-Kernel rows do not record CUDA, driver, or PyTorch versions, so their environments describe hardware only. A kernel is imported only once the semantics of its benchmark script have been reviewed.
- MLPerf serving rows report token throughput, which is the only quantity they measure. The TTFT and TPOT bounds shown beside them are thresholds set by the benchmark rules, and were not observed in the run.
- Hardware coverage reflects what the indexed sources have published and is not a systematic survey. The absence of a GPU or a kernel indicates only that no indexed source has published a result for it.
Data quality
A weekly job re-imports every source. An unexpected result stops that source before it writes, and an invariant check audits the whole catalog afterwards. The report is published at registry/reports/source-health.json; versioned catalog exports live under registry/exports. Errors can be reported from the action on any run page. Corrections retract or supersede a record and do not rewrite it.
Versions
Semantics change only through the publication of a new version. The current versions are manifests kernelindex.dev/v1alpha1 (schemas), ranking ranking-v1, deployability deployability-v2, and serving serving-v1. Each response states the version under which it was ranked, and each import records its parser version. Published runs and their digests do not change. The method history is recorded in the design doc's git log.
Privacy
KernelIndex records a small number of first-party counters, such as that a search occurred or that a result was opened. It sets no cookies and stores no identifiers, IP addresses, or query text. Counters are deleted after 90 days. Hosting adds cookieless, aggregate page-view counts (Vercel Web Analytics) with no persistent visitor identifier. Accounts store only the name and email address your sign-in provider shares, and you may delete your account at any time from /account. Full policy: Legal.