Preprint2026
How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
Venkata Naga Sai Vishnu Rohit Pulipaka, Anish Katta, Deva Rohit Reddy Peddireddy
This paper shows that a native C++ ONNX Runtime inference engine can score RLHF rewards faster than PyTorch eager mode and FastAPI on CPU, but torch.compile is faster on GPU, and batching strategy matters more than language or runtime choice.
0Jul 22, 2026
arXiv