3 ms·RL Is Bottlenecked by Inference. Scale It Independently11 points by alex000kim 2mo agodeleted 2mo ago[deleted]efiop 2mo agowhat did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.efiop 2mo agoah, nevermind. 3 engines seem cheaper overall too: 7x661s vs 5x1200s of allocated H100 time per step. Nice.