Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
intervieweratg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
intervieweratg
2y ago
Interesting, any reason to not use reasoning models? Is there anything 4o seems better at with respect to coding? I typically use o1 or o3-mini, but I am seeing that they just released an agent mode and, honestly, I think it depends on what
2.
▲
by
intervieweratg
2y ago
The results are in the paper and also in the announcement, I don’t think it’s too unusual. There is also an example of models cheating in SWE-Bench Verified in the appendix: ``` In response, o1 adds an underscore before filterable so that t