Skip to content
KernelIndex
Search⌘K

torch.compile (inductor)

PyTorch · python · MIT

Use it

Vendorable · source mirrored · MITView source →

No package. Vendor the mirrored source: 60 lines, MIT.

8_ResNetBasicBlock.py
curl "https://kernelindex.com/api/v1/implementations/kernelbench-l3-8-resnetbasicblock-torch-compile-inductor?include=source"
interfacepython · torch_compile_inductor
symbolModel.forward
Compatibility
measured onNVIDIA H100
declared hardwaredeclared only
architectures—
dtypes

Benchmark evidence

2 measurements across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
ResNetBasicBlockfp32 · [10, 3, 224, 224]
NVIDIA H100
875.0µs±7.84
#1 of 2
2026-03-05
ResNetBasicBlockfp32 · [10, 3, 224, 224]
NVIDIA H100
1.36ms±0.00
#1 of 2
2026-03-05

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:120cebdd3fdba1b873f379e190e6a967b9ecbbf0b188e518a0e6838a29437466
license declaredMIT
license concludedMIT
imported2026-08-26

Kernel source

8_ResNetBasicBlock.py60 lines
import torch
import torch.nn as nn
import torch.nn.functional as F

class Model(nn.Module):
    expansion = 1

    def __init__(self, in_channels, out_channels, stride=1):
        """
        :param in_channels: Number of input channels
        :param out_channels: Number of output channels
        :param stride: Stride for the first convolutional layer
        :param downsample: Downsample layer for the shortcut connection
        """
        super(Model, self).__init__()
        self.conv1 = nn.Conv2d(in_channels, out_channels, kernel_size=3, stride=stride, padding=1, bias=False)
        self.bn1 = nn.BatchNorm2d(out_channels)
        self.relu = nn.ReLU(inplace=True)
        self.conv2 = nn.Conv2d(out_channels, out_channels, kernel_size=3, stride=1, padding=1, bias=False)
        self.bn2 = nn.BatchNorm2d(out_channels)
        self.downsample = nn.Sequential(
            nn.Conv2d(in_channels, out_channels * self.expansion, kernel_size=1, stride=stride, bias=False),
            nn.BatchNorm2d(out_channels * self.expansion),
        )
        self.stride = stride

    def forward(self, x):
        """
        :param x: Input tensor, shape (batch_size, in_channels, height, width)
        :return: Output tensor, shape (batch_size, out_channels, height, width)
        """
        identity = x

        out = self.conv1(x)
        out = self.bn1(out)
        out = self.relu(out)

        out = self.conv2(out)
        out = self.bn2(out)

        if self.downsample is not None:
            identity = self.downsample(x)

        out += identity
        out = self.relu(out)

        return out
    
# Test code
in_channels = 3
out_channels = 64
stride = 1
batch_size = 10
num_classes = 1000

def get_inputs():
    return [torch.rand(batch_size, in_channels, 224, 224)]

def get_init_inputs():
    return [in_channels, out_channels, stride]
scrolls · 60 lines total

Source code from KernelBench, © 2023 Anne Ouyang, Simon Guo, Azalia Mirhoseini (Scaling Intelligence Lab, Stanford University), MIT License · MIT

Best evidence level for this revision: reported

JSON