submission 774929
x3C49 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 11 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-774929?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:bd9f1929680f17b7d2f3c274803af2c3de97d5586b222af846248d7c84640a13
license declaredunknown
license concludedunknown
authorsx3C49
imported2026-08-15
Kernel source
submission.py11 lines
import torch
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
input_tensor, output_tensor = data
# .to(float64).sum() allocates a full fp64 tensor in HBM then sums it:
# read fp32 -> write fp64 (new alloc) -> read fp64 -> scalar
# .sum(dtype=float64) fuses the cast into the reduction kernel:
# read fp32 -> scalar (no intermediate tensor)
# Returns a scalar (shape []) matching the reference's return type.
return input_tensor.sum(dtype=torch.float64).to(torch.float32)Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON