Fused linear + ORPO loss
ORPO preference loss over batch×seq×hidden activations projected through a vocab×hidden LM head against integer targets, computed chunked without materializing logits.
Reported evidence · last observed 2024-11-13. Reported by source; not independently reproduced.
pip install "liger-kernel==0.4.0"Reported evidence · last observed 2024-11-13. The source's designated baseline implementation. Reported by source; not independently reproduced.
Fastest by languagepython · Transformers · 12.9 ms · #1triton · Liger · 31.0 ms · 2.40×
Current records
Not measured on H100, B200 for this workload. Challenges →
Estimated floor 523.5 µs · record 24.66× above itestimate, not evidence ›
112.9ms1.00×Reported · Apache-2.0 · source2024-11-13stale
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
2LigerLiger-Kernel31.0ms2.40×Reported · BSD-2-Clause · source2024-11-13stale
2.40× slower than the baseline. Measured exactly what you asked. Reported by source; not independently reproduced.