Skip to content
KernelIndex
Search⌘K

Top k sampling from probs v151936

0 eligible runs
sampling

Top-k sampling from probabilities with vocab_size=151936. Keeps only the k highest probability tokens, renormalizes, then samples from the filtered distribution.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
Show all 9 implementations ›
flashinfer_wrapper_9c1e50faFlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
probsfp32 [batch_size, vocab_size]
top_kint32 [batch_size]
samplesint64 [batch_size]
Axes and behavior
batch_sizevariable
vocab_sizeconstant = 151936
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastop_k_sampling_from_probs_v151936modelqwen3-8bmodelqwen3-14bmodelqwen3-30b-a3bmodelqwen3-32b+1 modelssha256fb2439705ba9…
No source imports for this operation yet.How records are decidedJSON