Skip to content
KernelIndex
Search⌘K

Top k sampling from probs v129280

0 eligible runs
sampling

Top-k sampling from probabilities with vocab_size=129280. Keeps only the k highest probability tokens, renormalizes, then samples from the filtered distribution. Captured from DeepSeek V3.

Current records

No published measurement for the selected workload.

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
cuda
—
—
Apache-2.0 · source
triton
—
—
Apache-2.0 · source
Show all 9 implementations ›
flashinfer_wrapper_4ec4ec35FlashInfer-Bench baselines
python
—
—
Apache-2.0 · source

Semantics

Inputs and outputs
probsfp32 [batch_size, vocab_size]
top_kint32 [batch_size]
samplesint64 [batch_size]
Axes and behavior
batch_sizevariable
vocab_sizeconstant = 129280
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliastop_k_sampling_from_probs_v129280modeldeepseek-v3modeldeepseek-r1sha256cf78171e8ec4…
No source imports for this operation yet.How records are decidedJSON