Skip to content
KernelIndex
Search⌘K

Top k sampling from probs v128256

0 eligible runs
sampling

Top-k sampling from probabilities with vocab_size=128256. Keeps only the k highest probability tokens, renormalizes, then samples from the filtered distribution. Captured from Llama 3.1 8B.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
Show all 9 implementations ›
flashinfer_wrapper_d86b24bdFlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
probsfp32 [batch_size, vocab_size]
top_kint32 [batch_size]
samplesint64 [batch_size]
Axes and behavior
batch_sizevariable
vocab_sizeconstant = 128256
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastop_k_sampling_from_probs_v128256modelllama-3-1-8bmodelllama-3-1-70bmodelllama-3-1-405bmodelllama-3-2-3b+1 modelssha25650f0fc3c5df0…
No source imports for this operation yet.How records are decidedJSON