4 ms·
https://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/ Love those GPQA scores hovering around 5% when chance (on 4-way multi-choice) wou
by adt 1y ago
https://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
Love those GPQA scores hovering around 5% when chance (on 4-way multi-choice) would have got them 25%!
- gryfft 1y agoA stopped clock is right twice a day, but a running clock set to the wrong time is always wrong.
- parrit 1y agoThe RMS of wrongness of the running clock is probably lower.
- cwt137 1y agoNot always true! Your statement is only true when the running clock's speed is the same as time. Thus, regular time and the clock's time will never meet. If the clock is running faster than regular time, it will at point catch up to regular time and thus be correct for a split second. If the clock is slower than regular time, regular time will catch up to the clock and the clock will be right for a split second.
- actionfromafar 1y agoIf we are being pedantic, running clocks never run exactly the same as time. So they'll be right (very) much more seldom than the stopped clock, which is right twice a day.
- nathan_douglas 1y agoIf the clock is running backwards at very high speed, it would be right infinitely many times but the proportion of the time that it is right would approach some finite constant.
- k__ 1y agoMy girlfriend's microwave-clock runs faster than normal. Somehow this thing manages to accumulate an error of ~15 minutes in a month.
- patapong 1y agoAnd we haven't even touched on the issue of 24-hour format digital clocks, which can at most be right once per day if stopped!
- nthingtohide 1y ago> a running clock set to the wrong time is always wrong. Could be right within 15 min accuracy in the appropriate timezone. And such a mechanism can be corrected for in the postprocessing step.
- montebicyclelo 1y agoSo could do better than chance by excluding the option it's picked?
- dudeinhawaii 1y agoor.. A stopped clock is right twice a day; a mis-prompted LLM is wrong 19 times out of 20—but only because we handed it the wrong instruction sheet. Procedural error in testing perhaps? I'm not familiar with the methodology for GPQA.