Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
benjiro29
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
benjiro29
2mo ago
Ironically, we are also moving to more capable / faster models that use less power. DeepSeek V4 Flash 0731 is so extreme capable and comparability to a lot of models cheap to run. We are seeing stuff like AMD buying Taalas, with their
32.
▲
by
benjiro29
2mo ago
Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. What cost the most in API. Input, Cached Input, or Output. There you have your answer. Unfortunately, we have moved so much of the actual in
33.
▲
ROI of $100 Claude Code Subscription
(medium.com)
2 points
by
benjiro29
2mo ago
|
1 comments
34.
▲
by
benjiro29
2mo ago
A very interesting cost analyze of using AI and without by a software engineering. Interesting quote by OP on reddit: > So I wouldn't say "AI saved me $200k". What actually happened is that without it I'd have quit a
35.
▲
by
benjiro29
2mo ago
Because every company may have different needs that are not fulfilled by standards software. We have seen the large number of companies whose goal is to make custom software for other companies So is not always that reinventing is fun, but
36.
▲
by
benjiro29
2mo ago
> Claude code does some of this by handing off the "explore" agent work to haiku. That is not handing off to a specialized model, its just handing off to a lighter and interior model (compared to the parent model). That by itse
37.
▲
by
benjiro29
2mo ago
Repost from above: https://en.wikipedia.org/wiki/Film_industry#Largest_markets_ ... ("Largest markets by box office revenue") United States and Canada $8,870,000,000 2025[175] Europe
38.
▲
by
benjiro29
2mo ago
Whoever put those trailer together really knew how to sell it as a "do not bother to see movie". And frankly, Jason Momoa fatigue may also be a issue. Jason momoa playing Jason momoa, playing Jason momoa. Its like with The Rock pl
39.
▲
by
benjiro29
2mo ago
https://www.tbench.ai/leaderboard/terminal-bench/2.1 > 78.4 The real score is always the official benchmark. We need to see later if DS4 flash 0731 is going to maintain the score but we need to look at the offi
40.
▲
by
benjiro29
2mo ago
Probably the same. When the same base model is trained, the weight do not tend to change a lot. GLM 5.0 > 5.1 > 5.2 are the same base model, that just kept being trained. Weights hardly change as a result. Think in the like few percen
41.
▲
by
benjiro29
2mo ago
I think it does not matter. Most people who use these types of models are into the IT world, and will know/be informed very fast that there is a difference. People will likely also use the 0731 behind it. The only issue i see, is 3th p
42.
▲
by
benjiro29
2mo ago
MiMo, the overlooked sidekick to the hero. Will be interesting to see what Xiaomi bring to the table. These massive jumps in cheap models, is really great times!
43.
▲
by
benjiro29
2mo ago
But on Terminal bench, its * DS4 Flash: 82.7 * GPT 5.6 Luna: 75.7 For reference, that puts it on the third spot behind GPT 5.5 and Fable 5. For some reason GPT 5.6 Sol is not showing in the leaderboard. If it did, then DS4 Flash was number
44.
▲
by
benjiro29
2mo ago
Seems like it might be more advantageous to just adjust reasoning effort to retain cache. Some agents when you alter the reasoning level, partially or complete wipe the cache. Never assume that changing reasoning is a no-impact change. I
45.
▲
by
benjiro29
2mo ago
> You Could Have Come Up with ... Creating or combining to have something new, that does not already exist is actually freaking hard! The moment its presented and people go "o, that is not that difficult", "i was able to a
46.
▲
by
benjiro29
3mo ago
Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring o
47.
▲
by
benjiro29
3mo ago
I really do not understand how software developer think anymore. Using a LLM to translate a project in a short time, is by itself incredible. Just like one-shot whatever office clone. But what makes software is not the fast creation of a &q
48.
▲
by
benjiro29
3mo ago
Same answer i gave to somebody else up here... If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper. You see the issue? Its still a expensive
49.
▲
by
benjiro29
3mo ago
Then your comparing to a level of GPT 5.6 High, what is 50% cheaper then Opus Medium for the same intelligence / score. You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
50.
▲
by
benjiro29
3mo ago
> At half the price and less likely to auto-downgrade, it sounds like a reasonable claim Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8). Already posted this before,
51.
▲
by
benjiro29
3mo ago
Also the cost per task. https://www.vals.ai/benchmarks/vals_index !!! Vals !!! Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but
52.
▲
by
benjiro29
3mo ago
Not how it works ... Anything public viewable is still subject to a TOS. And the whole copyright or whatever still applies. If your argument was valid, we can scrape news websites and show their content freely. No ... You get sued and lose
53.
▲
by
benjiro29
3mo ago
> It was only 15 days. That, and usage was also half. Given the fact that we seen not any evidence beyond people "see, it one shot X game looks similar to Fable" (a lot of one shots look similar to other models). You expect to
54.
▲
by
benjiro29
3mo ago
People keep forgetting that over the last 6+ months a lot of increased action has been taken by OpenAI and Anthropic to detect and combat distillation. Several are public known. Combined with how short of a time Fable was around before K3 g
55.
▲
by
benjiro29
3mo ago
As somebody who did plenty of scraping for his own little projects, lets just say that reddit their security concerns are just PR for "we do not want to keep supporting old.reddit". While yes, you can not simply scrap new reddit a
56.
▲
by
benjiro29
3mo ago
!! Be careful when testing the model. A lot of people are testing it, and reporting disappointed results / benchmaxxxing claim. But do not realize that thinking has a issue with the default configuration. Important - make sure that TH
57.
▲
HuggingFace security incident: Guardrails vs. Open Models
(huggingface.co)
14 points
by
benjiro29
3mo ago
|
1 comments
58.
▲
by
benjiro29
3mo ago
The interesting aspect of this attack and the how guardrails prevented them from analyzing the attack. When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitt
59.
▲
by
benjiro29
3mo ago
The problem with these type of new / transcompile languages is not the langue, its the lacking libraries and 3th party assets. Great if just want the most basic hello world programs, but the moment you need ... for example a http serve
60.
▲
by
benjiro29
3mo ago
In regards to your post and the 16k reasoning output. Try setting reasoning levels yourself manually. We see in the benchmarks that one of the graphs shows low, mid, max, so its clearly there. I had the same issue with GLM 5.2 only offering
More ›