Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
XCSme
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
211.
▲
by
XCSme
4mo ago
Not loading for me, empty page (Brave/Windows)
212.
▲
by
XCSme
4mo ago
It is possible, just saying that I can't do it, and there's no logic behind it. You just have to be lucky and stumble upon a sequence that allows you to surive. The RNG is seeeded though, always same seed, so you can determinarica
213.
▲
by
XCSme
4mo ago
Can't go over 19, seems to be a lot of RNG, sometimes pieces spawn around protecting each-other and you stuck between them, so not much you can do.
214.
▲
by
XCSme
4mo ago
I'm just playing devil's advocate here. Yes, but still "how to cook" is not atomic. It involves knowing how to move stuff, how to measure, what "cooked" looks like in different environment (i.e. different light
215.
▲
by
XCSme
4mo ago
How would it know about Wikipedia and when to use it? From the tool description? If we had 100k such tools, then that wouldn't even fit in the context. This is only one example, plus if the topic is more complex, maybe it had to search
216.
▲
by
XCSme
4mo ago
Yup, you still need knowledge. Even if you have access to all the data and tools, you still need to know what to search for, what tools to use and to understand what the user is asking. Our computers can already do everything, have access t
217.
▲
by
XCSme
4mo ago
Check out my comparison too, it has some not-really-benchmarks too (between any two models actually, SVG generation test and CSS animation test): https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...
218.
▲
Show HN: One hundred LLMs Generating a HTML/CSS Solar System
(aibenchy.com)
5 points
by
XCSme
4mo ago
|
1 comments
219.
▲
MariaDB now has a DuckDB storage engine
(mariadb.org)
2 points
by
XCSme
4mo ago
|
0 comments
220.
▲
by
XCSme
4mo ago
Thanks for the feedback! What are you using Claude models for? Coding only? Computer use? Which harness?
221.
▲
by
XCSme
4mo ago
Also Claude/Fable models are quite bad at instructions following: https://artificialanalysis.ai/evaluations/ifbench
222.
▲
by
XCSme
4mo ago
On some it does yes, also in real usage. It avoided answering 2/21 tests in this specific benchmark mark, that's already 90% max score already.
223.
▲
by
XCSme
4mo ago
Well, most people were not liking Fable when it was available anyway, because it refused to answer questions very often.
224.
▲
by
XCSme
4mo ago
PS: Just added a cool feature, so you can filter the leaderboard for multiple models at once, by using a comma, like: https://aibenchy.com/?q=glm,claude
225.
▲
by
XCSme
4mo ago
I also tested it[0]: quite similar to GLM 5, a few percent better, 30% faster and 50% more expensive. [0]: https://aibenchy.com/?q=glm
226.
▲
by
XCSme
4mo ago
I think the problem is, as can also be seen on other benchmarks, is that most models nowadays are focused more and more purely on tool calling and coding. This means, that models are losing more and more general and domain-specific knowledg
227.
▲
by
XCSme
4mo ago
Oh, or you meant a smaller model than GLM-5.2 with similar capabilities?
228.
▲
by
XCSme
4mo ago
Which Opus? GLM-5.2 is already close to Opus-4.7 level: https://aibenchy.com/compare/anthropic-claude-opus-4-7-mediu...
229.
▲
by
XCSme
4mo ago
In my tests[0] GLM-5.2 is not much better than GLM-5, and overall DeepSeek V4 Flash seems to be the better/more cost-effective choice: [0]: https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high...
230.
▲
by
XCSme
4mo ago
Kind of, the 4-hour work workweek was one of the first book's I've read (I started late, was never interested), and it had many good insights that lead to me living a freer and more fulfilling life, that I might have not done othe
231.
▲
by
XCSme
4mo ago
So, are you saying that local models are maybe better than we give them credit? Because with some extra orchestration/processing we could improve the results?
232.
▲
by
XCSme
4mo ago
> You will notice multiple "thinking" parts per "turn" I thought that was the code harness simply minifying the outputs. Many models now no longer return the entire chain-of-thought (to avoid distillation attacks). So
233.
▲
by
XCSme
4mo ago
> The SOTA models are a deep orchestration of multiple models operating together it isn't a single mode I don't understand, why does it make you think this is the case? > how can GPT send thinking parts one after another wit
234.
▲
by
XCSme
4mo ago
Seems to be similar level to Kimi K.26, just that it's more token efficient and cheaper to run: https://aibenchy.com/compare/moonshotai-kimi-k2-6-medium/moo...
235.
▲
by
XCSme
4mo ago
Good point, everyone was expecting GPU prices to go down and all crypto bros to drop their mining rigs, but they can just repurpose them now...
236.
▲
by
XCSme
4mo ago
It also does A LOT better, for my hamster test: https://aibenchy.com/showcase/?q=claude#showcase=6efb87c28e3...
237.
▲
by
XCSme
4mo ago
Best hamster by far: https://aibenchy.com/showcase/?q=claude
238.
▲
by
XCSme
4mo ago
> We recently submitted a confidential S-1. We expect it to leak so we’re just announcing it. What?
239.
▲
by
XCSme
4mo ago
Not sure how I feel about Grok... [0] [0]: https://aibenchy.com/showcase/?q=grok
240.
▲
by
XCSme
4mo ago
I like how the Gemini 3.5 Flash (medium) one[0] added a ladder, so the hamster can get on the table. [0]: https://aibenchy.com/showcase/#showcase=eb4878dbff331c67
More ›