Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mindwraps
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
mindwraps
14d ago
Correct. But that's by design here. Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to
2.
▲
by
mindwraps
14d ago
One of the authors here - We had both run through a curated set of 64 long horizon coding tasks to evaluate cost/intelligence. Astra surprised us in this default setting. How is the rest of you faring (especially interested in those wh
3.
▲
Fable 5.1 vs. Astra for coding: Fable 2X more expensive per task but solves more
(aistack.imec-int.com)
2 points
by
mindwraps
14d ago
|
4 comments
4.
▲
Benchmarking Claude Code, Codex and Pi on SWE-Bench Pro: Same Accuracy, 2x Cost
(aistack.imec-int.com)
1 points
by
mindwraps
23d ago
|
0 comments
5.
▲
by
mindwraps
2mo ago
Solid point. We debated multiple optimization methods and decided to do a more vanilla run first for the setups tested. We’re moving on to some more typical optimization methods for our next article… As soon as our teams comes out a well de
6.
▲
by
mindwraps
2mo ago
Noted. Making sure we have a cleanly readable version asap. We wanted to stand out a bit vs. our mother-brand (imec), but fully agree some of us like a more sober reading experience. Feedback much appreciated.