5 ms·
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which directio
by intrepidkarthi 3mo ago
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions.