submission 153816
HayatoFujihara · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 99 lines, June 9 Researcher Reciprocity License v1.0.
grayscale_v2_6.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-153816?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3ca256a9e598cb3a8b55df96fe1bdff3bf6c8daecd1760d789630e6dc32e92bc
license declaredunknown
license concludedunknown
authorsHayatoFujihara
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
@triton.autotune(num-warps = 8
triton.Config({'BLOCK_SIZE': 4096}, num_warps=8, num_stages=4),stages = 4
triton.Config({'BLOCK_SIZE': 4096}, num_warps=8, num_stages=4),Kernel source
grayscale_v2_6.py99 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t
# -----------------------------------------------------------------------------
# Final Optimization: "Parallel Arithmetic & Lean Launch"
# -----------------------------------------------------------------------------
# 1. Parallel FMA: 積和演算の依存チェーンを断ち切り、R,G,Bの計算を同時に行わせます。
# 2. Optimized Launch: triton.cdiv などの関数呼び出しをやめ、直接計算します。
@triton.autotune(
configs=[
# A100の鉄板設定
triton.Config({'BLOCK_SIZE': 4096}, num_warps=8, num_stages=4),
triton.Config({'BLOCK_SIZE': 4096}, num_warps=8, num_stages=5), # 深いパイプライン
triton.Config({'BLOCK_SIZE': 2048}, num_warps=4, num_stages=4),
triton.Config({'BLOCK_SIZE': 2048}, num_warps=4, num_stages=5),
# 巨大ブロックも一応入れておく
triton.Config({'BLOCK_SIZE': 8192}, num_warps=8, num_stages=4),
],
key=['n_elements'],
)
@triton.jit
def grayscale_parallel_math_kernel(
input_ptr, output_ptr,
n_elements,
BLOCK_SIZE: tl.constexpr
):
pid = tl.program_id(axis=0)
# --- Input Block Pointer ---
in_ptr = tl.make_block_ptr(
base=input_ptr,
shape=(n_elements, 3),
strides=(3, 1),
offsets=(pid * BLOCK_SIZE, 0),
block_shape=(BLOCK_SIZE, 1),
order=(1, 0)
)
# --- Load (Issuing 3 loads back-to-back) ---
# ポインタを複製して advance させることで、独立したロード命令として発行
# コンパイラがスケジューリングしやすくなります
# R Channel
r = tl.load(in_ptr, boundary_check=(0,), padding_option="zero")
# G Channel
in_ptr_g = tl.advance(in_ptr, (0, 1))
g = tl.load(in_ptr_g, boundary_check=(0,), padding_option="zero")
# B Channel
in_ptr_b = tl.advance(in_ptr_g, (0, 1))
b = tl.load(in_ptr_b, boundary_check=(0,), padding_option="zero")
# --- Parallel Computation (ILP) ---
# 以前: gray = r*c1; gray += g*c2; gray += b*c3 (直列依存)
# 今回: term1, term2, term3 を並列計算 -> 最後に合算
term1 = r * 0.2989
term2 = g * 0.5870
term3 = b * 0.1140
# 合算
gray = term1 + term2 + term3
# --- Output Block Pointer & Store ---
out_ptr = tl.make_block_ptr(
base=output_ptr,
shape=(n_elements,),
strides=(1,),
offsets=(pid * BLOCK_SIZE,),
block_shape=(BLOCK_SIZE,),
order=(0,)
)
gray_flat = tl.reshape(gray, (BLOCK_SIZE,))
tl.store(out_ptr, gray_flat, boundary_check=(0,))
def custom_kernel(data: input_t) -> output_t:
input_tensor, output_tensor = data
n_elements = output_tensor.numel()
# Python Overhead Optimization:
# triton.cdiv(n, b) は (n + b - 1) // b と等価ですが、
# 関数呼び出しオーバーヘッドを嫌って直接書きます。
# ベストな BLOCK_SIZE は autotune で選ばれますが、
# 起動時にはその値を使ってグリッドを計算する必要があります。
grid = lambda META: ((n_elements + META['BLOCK_SIZE'] - 1) // META['BLOCK_SIZE'], )
grayscale_parallel_math_kernel[grid](
input_tensor, output_tensor,
n_elements
)
return output_tensorscrolls · 99 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON