Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
re-thc
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
re-thc
13d ago
There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
2.
▲
by
re-thc
14d ago
> You refuse a direct order. I never said. It's not about outright refuse. It's how to deliver better outcomes. It's not a win or lose situation. Don't treat it like that. > If your boss says that he does not care
3.
▲
by
re-thc
15d ago
> Where LLMs excel is in code-level bugs (as opposed to system bugs, design bugs, architecture bugs, integration bugs, etc). Blame the benchmarks game. They're optimizing for that and that's what those things are measuring.
4.
▲
by
re-thc
15d ago
> Work for business people who want fast results. Agentic coding gets you to something presentable much faster at the cost of code quality. I have never seen a customer or business person care about that. That's always false. It
5.
▲
by
re-thc
21d ago
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. It's not useful because the benchmarks often measure the wrong thing. They're here yapping about
6.
▲
by
re-thc
21d ago
It did used to use Haiku but that model is now too too far behind…
7.
▲
by
re-thc
23d ago
Hardware definitely has longer lifecycle than AI model releases at this point. You don't see Nvidia and AMD fighting every other month over the latest cards.
8.
▲
by
re-thc
23d ago
Overtaken in cost per token.
9.
▲
by
re-thc
24d ago
> DeepSWE is a very big deal It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".
10.
▲
by
re-thc
24d ago
> 4.0 Flash we will finally get Gemini 3.5 Pro Nah, we'll just get the 4.0 Pro Preview.
11.
▲
by
re-thc
24d ago
> Beginning to think Google is a dark horse in this race Google was so hyped up early Gemini 3 era (only some months ago). And now dark horse? The TPU takeover almost crashed nvidia and everyone else.
12.
▲
by
re-thc
25d ago
> people are still saying the Codex limits are more generous. They're not They are if you follow Tibo on the resets.
13.
▲
by
re-thc
25d ago
With OpenAI you can also apply for the security program, which doesn't require you to be a certified pentester (as per Anthropic).
14.
▲
by
re-thc
25d ago
> IMO, Codex is worse than Claude with Fable. Fable easily trips its safe guards. You can be 95% complete with the plan for it to trip and then lose it all. Anything is better than nothing.
15.
▲
by
re-thc
25d ago
That's load bearing!
16.
▲
by
re-thc
25d ago
The biggest change is the price cut of course.
17.
▲
by
re-thc
28d ago
Agent scale!
18.
▲
by
re-thc
28d ago
> Otherwise the premise for coding using LLMs is essentially untrue. Why? Written by does not mean designed by etc. There's a lot more to it.
19.
▲
by
re-thc
29d ago
Go has API pricing + this weird scaling of how much is it worth. Some models get $60 of usage, some $30 and some $15 etc.
20.
▲
by
re-thc
29d ago
> that would prove rather embarrassing for Anthropic Not really, in that you just work with different constraints. Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and Chinese labs in general
21.
▲
by
re-thc
1mo ago
And it could have expanded elsewhere?
22.
▲
by
re-thc
1mo ago
> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts It's what people know. Opus is just the common target.
23.
▲
by
re-thc
1mo ago
> The export controls were revoked before Zai is on another "export control" list outside the broader 1. Doesn't help.
24.
▲
by
re-thc
1mo ago
Which 9/10 times hasn't been great anyway (stock reaction).
25.
▲
by
re-thc
1mo ago
> From a biased source, but would be big if true. I've had great results with GLM 5.2. It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
26.
▲
by
re-thc
1mo ago
2. There was a new checkpoint. Official.
27.
▲
by
re-thc
1mo ago
With vision on top
28.
▲
by
re-thc
1mo ago
the outperform Fable was a mid (not completed) benchmark run. Real results were lower.
29.
▲
by
re-thc
1mo ago
> _China doesn't need to invade Taiwan._ All your points consider what China can do but not how others would respond. Blocking Taiwan when the whole AI economy is running on it will invoke lots of countries to go up in arms. There w
30.
▲
by
re-thc
1mo ago
> this model is around 64B and can run on laptops. We're lucky if it'd fit in 1 DGX Spark. Laptops - nah, unless you mean like an M5 Max with 128GB of RAM then maybe. > Thats the reason behind the hype. The hype is imagine D
More ›