Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
achalpandey
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
achalpandey
3mo ago
Hmm, I am trying to benchmark cost/quality for real world sessions. In that scenario "model resourcefulness" and efficiency is actually a good thing. Why spend tokens working through a solution when you can simply look it up?
2.
▲
by
achalpandey
3mo ago
UPDATE: turns out "some" models know how to game the premise. They simply lookup the solution to the exact solved SWE bench problems! Haha Because I don't want to impose artificial constraints like no network access, I'm
3.
▲
by
achalpandey
3mo ago
Here's my current plan, the "session" will be made up of multiple SWE bench tasks stitched together. Each "task" is the equivalent of a new user query and we also pre-program "cache expiration" (sleep for
4.
▲
by
achalpandey
3mo ago
Thank you! Both of those papers are super new and super relevant. The only big gap left is that they aren't using claude code/codex as harnesses. I'll try to reuse their constructed user sessions. PS: Your work at stet is als
5.
▲
by
achalpandey
3mo ago
Looking for feedback and thoughts. Here's a link to my one-page spec: https://docs.google.com/document/d/e/2PACX-1vRu5Fv5-KTJDnCEx...
6.
▲
How are you measuring Claude Code and Codex performance?
3 points
by
achalpandey
3mo ago
|
9 comments
7.
▲
by
achalpandey
11mo ago
Hey HN, My co-founder and I built this because we were frustrated. Typing our ideas was too slow. Voice notes were even worse, "great for capture, but useless for action". They just became this "messy, unusable inbox". I
8.
▲
Show HN: Speak your mind and get a prioritized action plan instantly
(apps.apple.com)
1 points
by
achalpandey
11mo ago
|
1 comments