Skip to content
KernelIndex
Search⌘K

Top p sampling from probs v151936

0 eligible runs
sampling

Top-p (nucleus) sampling from probabilities with vocab_size=151936. Filters probabilities using cumulative probability threshold, then samples from the filtered distribution. Captured from Qwen 3 30B A3B.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
flashinfer_wrapper_32ca24afFlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
probsfp32 [batch_size, vocab_size]
top_pfp32 [batch_size]
samplesint64 [batch_size]
Axes and behavior
batch_sizevariable
vocab_sizeconstant = 151936
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastop_p_sampling_from_probs_v151936modelqwen3-8bmodelqwen3-14bmodelqwen3-30b-a3bmodelqwen3-32b+1 modelssha25689a950a49d1c…
No source imports for this operation yet.How records are decidedJSON