Skip to content
KernelIndex
Search⌘K

PyTorch

Chunked distillation JSD loss · vocab = 128256 · hidden = 4096 · tokens = 4096 · bf16 · run 01a0202b-2178…
●passedReported evidence

Not re-observed recently.

Primary measurement
107.1msmedian

Rank 2 in its comparison group · source-native comparison · observed 2024-12-03

Compare with #1 →

ms 20 107.08905792236328 · ms 50 107.08905792236328 · ms 80 107.08905792236328

Identity
implementationPyTorch
projectPyTorch
revisionunknown
workloadvocab = 128256 · hidden = 4096 · tokens = 4096 · bf16
comparison keysha256:8bb943d36a5f8fa1…
sourceLiger-Kernel benchmarks
external iddistill_jsd_loss/full/torch/nvidia-h100-80gb-hbm3/hidden4096-tokens4096-vocab128256-bf16/2024-12-03-08-00-25
sha256:2d3a5dbf87432863aa2268f…

Correctness

Marked passed by the source; the correctness policy was not published.

Workload

vocab128256
hidden4096
tokens4096
targetint64 [4096]
student_inputbf16 [4096, 2048]
teacher_inputbf16 [4096, 4096]
student_weightbf16 [128256, 2048]
teacher_weightbf16 [128256, 4096]
definition comparatornot_asserted

Measurements

latency · median107.1 ms
latency · p20107.1 ms
latency · p80107.1 ms

Protocol

harnessLiger-Kernel benchmark scripts
timercuda_events
primaryStatisticmedian
comparabilityFamilyliger_kernel_bench

Environment

gpuNVIDIA H100 (sm_90)

Artifacts

No artifacts published with this run.

Replications and notes

No attestations yet.

Community attestations. They never change the evidence level; only a KernelIndex-controlled rerun does.

Add a reproduction or note
I

Published under your account name. Community attestations never change the evidence level; a maintainer can hide abuse.

Canonical manifest

Show manifest
{
  "run": {
    "kind": "BenchmarkRun",
    "spec": {
      "status": "passed",
      "timing": {
        "latencyNs": {
          "median": 107089058
        },
        "primaryStatistic": "median"
      },
      "observedAt": "2024-12-03T08:00:25.000Z",
      "measurements": [
        {
          "unit": "ns",
          "value": 107089058,
          "metric": "latency",
          "statistic": "p20"
        },
        {
          "unit": "ns",
          "value": 107089058,
          "metric": "latency",
          "statistic": "p80"
        }
      ],
      "sourceNative": {
        "source": "liger-kernel-bench",
        "metrics": {
          "ms_20": 107.08905792236328,
          "ms_50": 107.08905792236328,
          "ms_80": 107.08905792236328
        },
        "benchmark": "distill_jsd_loss/full",
        "externalId": "distill_jsd_loss/full/torch/nvidia-h100-80gb-hbm3/hidden4096-tokens4096-vocab128256-bf16/2024-12-03-08-00-25"
      },
      "protocolDigest": "sha256:5b02ee9a2140fbc6fcc1c89800f6c177ba22718c8897cb532bf0ba91867c1e58",
      "workloadDigest": "sha256:4691706ccbdb13fdc9cff83f292abb6c785229fa153870886bc0c4b75cedd0fe",
      "environmentDigest": "sha256:2750ae381d6de582f1eb6b9008e0987463d19ede8af127d97d76c5ae09121393",
      "implementationDigest": "sha256:7f1248e6e0aaddc0af487dbc9d280632c701b28f9bce8b6a3bb19e06bc2cd48d"
    },
    "metadata": {
      "name": "liger-distill-jsd-loss-full-torch-nvidia-h100-80gb-hbm3-hidden4096-tokens4096-vocab128256-bf16-2024-",
      "title": "Chunked distillation JSD loss · torch · full",
      "labels": {
        "liger_version": "0.4.2"
      }
    },
    "apiVersion": "kernelindex.dev/v1alpha1"
  },
  "protocol": {
    "kind": "BenchmarkProtocol",
    "spec": {
      "harness": {
        "name": "Liger-Kernel benchmark scripts",
        "repository": "https://github.com/linkedin/Liger-Kernel"
      },
      "measurement": {
        "timer": "cuda_events",
        "primaryStatistic": "median"
      },
      "comparability": {
        "notes": "Timed: the forward pass plus the backward pass, via triton.testing.do_bench with quantiles 0.5/0.2/0.8 (median with p20/p80 spread). The upstream CSV records no CUDA, driver, or torch version; the Liger release rides each run's labels. Comparable only within one kernel, workload, timed pass, and GPU.",
        "family": "liger_kernel_bench"
      }
    },
    "metadata": {
      "name": "liger-bench-do-bench-full",
      "title": "Liger benchmark harness · full pass"
    },
    "apiVersion": "kernelindex.dev/v1alpha1"
  },
  "environment": {
    "kind": "ExecutionEnvironment",
    "spec": {
      "hardware": {
        "vendor": "nvidia",
        "product": "NVIDIA H100",
        "formFactor": "80GB HBM3",
        "architecture": "sm_90"
      },
      "software": {}
    },
    "metadata": {
      "name": "liger-bench-nvidia-h100-80gb-hbm3",
      "title": "Liger benchmark host · H100 80GB HBM3"
    },
    "apiVersion": "kernelindex.dev/v1alpha1"
  }
}
Cite this record (permalink, digest, access date)
Report an issue with this run

A maintainer reviews every report. If accepted, the record is retracted or superseded; its history stays visible.