4 ms·
Suspect there a lot of methodology flaws. Depends a lot on codebase context but the first line of the exposed prompt says "Monday" - which monday, UTC? Unless a
by evalmaster123 13d ago
Suspect there a lot of methodology flaws. Depends a lot on codebase context but the first line of the exposed prompt says "Monday" - which monday, UTC? Unless a small task set is exposed we cannot be sure if this variance is due to benchmark or models. Also very surprising to see GLM/ other models perform better on few tasks.
Currently this is a big TRUST ME BRO