Fused linear + KTO loss
KTO preference loss: policy and reference batch×seq×hidden activations project through separate biased vocab×hidden LM heads against integer targets, with per-sequence boolean preference labels and a precomputed KL term (beta 0.1).
Reported evidence · last observed 2025-03-03. Reported by source; not independently reproduced.
pip install "liger-kernel==0.5.4"Reported evidence · last observed 2025-03-03. The source's designated baseline implementation. Reported by source; not independently reproduced.
Fastest by languagepython · Transformers · 3.88 ms · #1triton · Liger · 4.00 ms · 1.03×
Current records
Not measured on B200 for this workload. Challenges →
Estimated floor 158.2 µs · record 24.5× above itestimate, not evidence ›
13.88ms1.00×Reported · Apache-2.0 · source2025-03-03stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
2LigerLiger-Kernel4.00ms1.03×Reported · BSD-2-Clause · source2025-03-03stale
1.03× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.