Best reported throughput per model and scenario. Open a row for every configuration, the Pareto frontier, and exclusions. MLPerf measures token throughput only; the TTFT and TPOT limits above compare against each scenario's declared rules, never against latency KernelIndex has not seen measured.
Model
Benchmark
Best tokens/s
Stack
Hardware
Configs