3 ms·
Reframe this as scaling test time compute using a human in the loop as the reward model. o1 is effectively trying to take a pass at automating that manual effo
by deepsquirrelnet 2y ago
Reframe this as scaling test time compute using a human in the loop as the reward model.
o1 is effectively trying to take a pass at automating that manual effort.