Skip to content
KernelIndex
Search⌘K

Top p sampling from probs v128256

0 eligible runs
sampling

Top-p (nucleus) sampling from probabilities with vocab_size=128256. Filters probabilities using cumulative probability threshold, then samples from the filtered distribution.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
flashinfer_wrapper_5df4fa0bFlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
probsfp32 [batch_size, vocab_size]
top_pfp32 [batch_size]
samplesint64 [batch_size]
Axes and behavior
batch_sizevariable
vocab_sizeconstant = 128256
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastop_p_sampling_from_probs_v128256modelllama-3-1-8bmodelllama-3-1-70bmodelllama-3-1-405bmodelllama-3-2-3b+1 modelssha2566dc343bc35c6…
No source imports for this operation yet.How records are decidedJSON