Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joshmlewis
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
1.
▲
by
joshmlewis
5mo ago
It's interesting they use output tokens as an eval because all tokens are not made equal. Even from model to model (like Opus 4.6 to Opus 4.7) the tokenizer can be different and it's no longer an apples to apples comparison. No on
2.
▲
by
joshmlewis
8mo ago
Speechify has been good for me although there might be better / cheaper alternatives I'm not aware of.
3.
▲
by
joshmlewis
8mo ago
I think the OP was implying that it's probably already baked into its training data. No need to search the web for that.
4.
▲
by
joshmlewis
8mo ago
"They" being the guy (Peter Steinberger) who created it as a personal project that he open sourced.
5.
▲
by
joshmlewis
9mo ago
This is cool but as someone that's built an enterprise grade agentic loop in-house that's processing a billion plus tokens a month, there are so many little things you have to account for that greatly magnify complexity in real wo
6.
▲
by
joshmlewis
1y ago
This just feels like the whole complicated TODO workflows and MCP servers that were the hot thing for awhile. I really don't believe this level of abstraction and detailed workflows are where things are headed.
7.
▲
by
joshmlewis
1y ago
This should not really be necessary and is more of a workaround for bad patterns / prompting in my opinion.
8.
▲
by
joshmlewis
1y ago
How big is your claude.md file? I see people complain about this but I have only seen it happen in projects with very long/complex or insufficient claude.md files. I put a lot of time into crafting that file by hand for each project be
9.
▲
by
joshmlewis
1y ago
It's also not a coincidence that Slack is neutering the ability to access channel history via the API very soon. With a very generous rate limit of 2 requests per minute I believe it was and a max of ~10 messages. This is already enfor
10.
▲
by
joshmlewis
1y ago
One of the biggest nuggets people need to take away from this: > At one point we tried “improving” the prompt with Claude’s help. It ballooned to 1,500 words. The agent immediately got slower and dumber. We went back to 103 words and it
11.
▲
by
joshmlewis
1y ago
As someone who builds AI products and having used agentic coding tools since they came out (often with Rails projects), I don't get this. There was a similar project called Rails MCP Server which said: > "This Rails MCP Server
12.
▲
by
joshmlewis
1y ago
It is funny how it can be like this sometimes. I think a lot depends on coding styles, languages, prompting, etc.
13.
▲
by
joshmlewis
1y ago
Cursor
14.
▲
GPT-5 really likes tool calling
(promptslice.com)
2 points
by
joshmlewis
1y ago
|
0 comments
15.
▲
by
joshmlewis
1y ago
When it came out on Tuesday I wanted to throw my laptop out of the window. I don't know what happened but results were total garbage earlier this week. It got better the past couple days but so far with gpt-5 being able to solve proble
16.
▲
by
joshmlewis
1y ago
Whoosh, it went right over my head.
17.
▲
by
joshmlewis
1y ago
The data is made up, the point is to see how models respond to the same input / scenario. You're able to create whatever tools you want and import real data or it'll generate fake tool responses for you based on the prompt an
18.
▲
by
joshmlewis
1y ago
I would highly doubt it. Even when you BYOK inside of Cursor they still say it's routed through their servers.
19.
▲
by
joshmlewis
1y ago
I noticed it was taking awhile on the first large-ish task I gave it. I'm assuming it was just a bit overloaded at the moment.
20.
▲
by
joshmlewis
1y ago
Where'd you get 720 from?
21.
▲
by
joshmlewis
1y ago
Did I say GPT-5? I said o3. :) That was a rebuttal to you saying you have never needed to add your key to use an OpenAI model before.
22.
▲
by
joshmlewis
1y ago
It seems to be trained to use tools effectively to gather context. In this example against 4.1 and o3 it used 6 in the first turn in a pretty cool way (fetching different categories that could be relevant). Token use increases with that kin
23.
▲
by
joshmlewis
1y ago
It's free in Cursor for the next few days, you should go try it out if you haven't. I've been an agentic coding power user since the day it came out across several IDE's/CLI tools and Cursor + GPT-5 seems to be a gr
24.
▲
by
joshmlewis
1y ago
It does seem to be doing well compared to Opus 4.1 in my testing the last few hours. I've been on the Claude Code 200 plan for a few months and I've been really frustrated with it's output as of late. GPT-5 seems to be a step
25.
▲
by
joshmlewis
1y ago
It does really well at using tool calls to gain as much context as it can to provide thoughtful answers. In this example it did 6! tool calls in the first response while 4.1 did 3 and o3 did one at a time. https://promptslice.com
26.
▲
by
joshmlewis
1y ago
I've been testing it against Opus 4.1 the last few hours and it has done better and solved problems Claude kept failing at. I would say it's definitely better, at least so far.
27.
▲
by
joshmlewis
1y ago
You had to use your own key for o3 at least. > Note that BYOK is required for this model. Set up here: https://openrouter.ai/settings/integrations https://openrouter.ai/api/v1/models
28.
▲
by
joshmlewis
1y ago
It's more efficient with tools for one and the input cost is cheaper (which is where a lot of the cost is). See comparison between GPT-5, 4.1, and o3 tool calling here: https://promptslice.com/share/b-2ap_rfjeJgIQs
29.
▲
by
joshmlewis
1y ago
I am convinced. I've been giving it tasks the past couple hours that Opus 4.1 was failing on and it not only did them but cleaned up the mess Opus made. It's the real deal.
30.
▲
by
joshmlewis
1y ago
It's a really good model from my testing so far. You can see the difference in how it tries to use tools to the greatest extent when answering a question, especially compared to 4.1 and o3. In this example it used 6! tool calls in the
More ›