Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dudeinhawaii
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
dudeinhawaii
25d ago
Your post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53. That then made me realize that they lower the bars of tied scores so on the site it looks like Astr
2.
▲
by
dudeinhawaii
25d ago
I have not experienced this (yet) but I have with models in the past. I think it's important to have a solid benchmark where you KNOW there's a difference in model performance. I have one around 3D modeling that models really land
3.
▲
by
dudeinhawaii
26d ago
Great site, triggered memories! haha. To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen is "overly helpful". Models nowadays want to double-triple-quadrup
4.
▲
by
dudeinhawaii
1mo ago
I don't think this is quite true. We have other examples. Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while Fable
5.
▲
by
dudeinhawaii
1mo ago
I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".
6.
▲
by
dudeinhawaii
1mo ago
Honestly, it's platform dependent and "OK" at best, "Mediocre" at worst (Agy on Windows). Gemini is great via the Chat interface and decent via Github Copilot. I honestly hate it via Antigravity CLI because their sa
7.
▲
by
dudeinhawaii
2mo ago
You are re-compressing information that is in-effect meaningless because it's all decompression artifacts. The AI had a nugget of data and decompressed that into a flood of text. The exhausting thing is that we're then trying to r
8.
▲
by
dudeinhawaii
2mo ago
It would make reading and comparing a bit easier if the data was sorted by a dimension.
9.
▲
by
dudeinhawaii
2mo ago
That's a good point and something I encountered yesterday. On a multi-agent task, Gemini was the only model that got near the end, ran tests, saw it had issues, took a screenshot, saw the issues, and then said, "I'll mark it
10.
▲
by
dudeinhawaii
2mo ago
I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but they never bothere
11.
▲
by
dudeinhawaii
2mo ago
My experience has been a difference between "applied intelligence" and "breadth of intelligence". Fable is the theoretical computer scientist while Opus is the Staff engineer who will implement it. I find that Opus has c
12.
▲
by
dudeinhawaii
2mo ago
Grok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way. It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with dimi
13.
▲
by
dudeinhawaii
3mo ago
I think most of the comments in this thread are missing the point. It's not about whether it's legal/ethical/etc. It's about the narrative that "Chinese models are at Fable level". The truth (if correct) i
14.
▲
by
dudeinhawaii
3mo ago
You should really provide more on your methodology because as it stands, it really doesn't pass the sniff test. GPT-5.6 Sol on Low beats Fable Medium by 10% and Gemini-3.6 Flash then beats them both? Fable is number 20? This does not m
15.
▲
by
dudeinhawaii
3mo ago
This is a perpetual pet peeve of mine with LLMs. They will always opt to ensure "backwards compatibility" with a codebase built 30 seconds ago. I always have to explicitly state, "we are making a clean break to v1.0, do not i
16.
▲
by
dudeinhawaii
3mo ago
In benchmarks for a product I'm working on I've noticed that Sol is hard to "contain". It will _always_ find the most effective way to game the system and dramatically outperform all other models. Fable 5 isn't an a
17.
▲
by
dudeinhawaii
3mo ago
I use all of the major providers daily and I tend to go to Gemini for "fast lookups" where a good enough answer is probably OK. I use ChatGPT and Claude for anything where it matters and generally when I invoke all three -- Gemini
18.
▲
by
dudeinhawaii
3mo ago
"We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue." Specific and helpful so that's good.
19.
▲
by
dudeinhawaii
3mo ago
Anthropic cried "I yield" in the "reset quota" wars. Someone has to pay for all those tokens. Hopefully that's not the case though, I was retaining a buffer to hit it hard this weekend.
20.
▲
by
dudeinhawaii
3mo ago
Watching the videos and reading the methodology, it feels like "scientist one-shot music video with LLMs" which is... useful but in no way represents how one would use the models. If nothing else, it serves as a great reminder tha
21.
▲
by
dudeinhawaii
3mo ago
Weird question that popped into my mind (not a judgement on this), but is there a similar jump in prosecutions for vehicular manslaughter or are these "whoopsie'd" away? Seeing today's distracted drivers, driving their m
22.
▲
by
dudeinhawaii
4mo ago
I don't want to pile onto the conspiracy thinking but I was just wondering how Anthropic was going to foot the bill for the clearly larger and more expensive model being run on millions of Claude Code subscriptions, subsidized until th
23.
▲
by
dudeinhawaii
4mo ago
This was one of the more amusing things I noticed very early on. I (and countless others) used AI to write war sims. The second I added nuclear silo construction; the next run was instantly nuclear Armageddon. One could argue that the LLMs
24.
▲
by
dudeinhawaii
4mo ago
No offense but your responses sound like AI or engagement farming. I think the "why are these rules good" is self-evident to anyone who read the comment.
25.
▲
by
dudeinhawaii
4mo ago
On the one hand, this feels very pragmatic. On the other hand, it feels like what people who weren't great software engineers say. It's kind of a craft. I can't imagine an exceptional artist saying "screw the craft, I ju
26.
▲
by
dudeinhawaii
4mo ago
This is the first time I saw a model pop-up on HN and didn't really care. Model exhaustion? It looks interesting but not exciting. While I'd normally _love_ incremental improvements --- I think the recent ones are far too minor to
27.
▲
by
dudeinhawaii
5mo ago
I'm not an Alexa user myself but I have watched my wife interact with it for around 5years now. The new Alexa powered by an LLM is objectively better that previous Alexa in a few ways. This much was apparently from day one and has only
28.
▲
by
dudeinhawaii
7mo ago
This has been my approach and of course what you lose is the "random and surprising" (maybe good) but also the "evolutionary" aspect. So, if you write strong tooling (even with AI) around the connection points - you can
29.
▲
by
dudeinhawaii
7mo ago
Why would I want non-deterministic behavior here though? If I want to max uptime, I write a tool to track/monitor. Then write a small agent (non-ai) that monitors those outputs and performs your remediation actions (reset something, cl
30.
▲
by
dudeinhawaii
7mo ago
It's not, hence the "don't post AI slop as your comment" posting a few days back that had 1000+ comments. Currently an unsolved problem - just stealthier on some platforms than others. Trigger the right topic on HN and t
More ›