Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tarruda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
tarruda
2mo ago
One of the best things about this version is that it is trained in the codex harness. It feels just as good as OpenAI models in using codex tools, but extremely cheap and with 1M context
32.
▲
by
tarruda
3mo ago
It is not that they don't release open weights, but some users report that they are significantly inferior to the closed versions.
33.
▲
by
tarruda
3mo ago
> You can't. Not if you're in the minority Is that a bad thing?
34.
▲
by
tarruda
3mo ago
> It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved. I suspect the value is in showing the potential that LLMs have in developing new breakthro
35.
▲
by
tarruda
3mo ago
Assuming steady 1 tok/second generation (which seems to be the case for M5 Max macbook), wait 1 day for a 86400 token response. In some configurations it can be as slow as 0.1 tok/s, so be prepared to wait for 10 days.
36.
▲
by
tarruda
3mo ago
Recently tried the pelican test on GPT-OSS which was probably one of the best local models of 2025. So cool to see how models have improved in the SVG pelican!
37.
▲
by
tarruda
3mo ago
> given they are pretty close in size One thing that might not be obvious about about DSV4 is how much innovation the Deepseek team implemented in its architecture. When llama.cpp fully supports its lightning indexer, the full 1M conte
38.
▲
by
tarruda
3mo ago
Hopefully this distillation will lead Alibaba to release more powerful open weights LLMs, contributing to the democratization of AI.
39.
▲
by
tarruda
4mo ago
Vibe thinker also beats Opus 4.5
40.
▲
by
tarruda
4mo ago
If your framework desktop is the 128G Strix Halo, I recommend giving Qwen 3.5 122B-A10B a shot. This Q5_K_M quant should be near lossless and fit with full 256K context in about 100GB of RAM: https://huggingface.co/AesSedai&
41.
▲
by
tarruda
4mo ago
I don't feel like AI coding has ruined my skills, and I could go back to manual coding any time. However, I cannot build a good mental model of a software component that I didn't write myself, and that can affect future maintenanc
42.
▲
by
tarruda
4mo ago
What I find fascinating is the idea that there might be a set of "secret" tweaks that when applied to those weights (or even smaller models) could result in an intelligence simulation that could vastly surpass even something like
43.
▲
by
tarruda
4mo ago
Not as much as Qwen, since apparently 3.6 35B surpassed Opus 4.7 https://x.com/simonw/status/2044830134885306701
44.
▲
by
tarruda
4mo ago
I don't think there's any incentive for Nvidia to make this a Windows-only device, so most likely it will be fully supported on Linux, just like their GPUs are.
45.
▲
by
tarruda
4mo ago
> This also means that, according to our plans, Zig will have to propagate "stackless-ness" upwards in the call chain while analyzing the code (thus making Future.await not special per-se). Very interesting, will be following Z
46.
▲
by
tarruda
4mo ago
> there's an accepted proposal to bring them back, in which case any function that calls await, or that otherwise has a suspension point, would have to be transformed into a stackless coroutine by the compiler, yes. The plan is for
47.
▲
by
tarruda
4mo ago
Fixed it. thanks!
48.
▲
by
tarruda
4mo ago
> especially with the new IO mechanism which allows supper efficient code that looks good whether it's implemented single-threaded, multi-threaded or just via an event loop! I had some trouble understanding how the async/await
49.
▲
by
tarruda
4mo ago
The official Q4_K_S gguf is quite good and has very good 35 tps generation on a M1 mac studio. Should be much faster on recent Macs, especially M5.
50.
▲
Step 3.7 Flash
(static.stepfun.com)
48 points
by
tarruda
4mo ago
|
16 comments
51.
▲
by
tarruda
4mo ago
> One of the most prominent improvements in Opus 4.8 is its honesty. Does that mean it no longer deletes or changes tests to make it pass?
52.
▲
by
tarruda
5mo ago
> safer bet as a dependency. The recent 1 million line vibe coded PR suggests it is not so reliable as a dependency.
53.
▲
by
tarruda
5mo ago
> That's impressive getting a 397B down to <110GB It is higher than 110GB. MacOS allows up to 125G of the RAM to be shared with GPU, so it is certainly less than that! > HF link is broken though! Doesn't seem broken to me
54.
▲
by
tarruda
5mo ago
> I'm questioning ROI If by ROI you mean saving more money than using paid APIs, then I don't think it is worth it. All you gain is full sovereignty over your AI usage.
55.
▲
by
tarruda
5mo ago
> 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes. I run 2.54 BPW 397B Qwen 3.5 GGUF on a 128G mac studio at 20 tokens/second generation and 200 tokens/second pro
56.
▲
by
tarruda
5mo ago
I only tried a very early version of that when it was just a llama.cpp fork and Qwen was certainly better in my tests. But I was not super impressed with deepseek 4 flash using it from the official API either, so it doesn't seem quanti
57.
▲
by
tarruda
5mo ago
> What’s the price point for getting into that sweet spot? In October/2024 I got my Mac studio M1 ultra with 128G, IIRC it was ~$2500. With recent prices explosion, it has certainly gotten more expensive. https://frame.wo
58.
▲
by
tarruda
5mo ago
I have a 128G mac studio and even 397B was a happy surprise to me due to its high quantization resilience. I've created a 2.54BPW quant that fit on my hardware with 128k context, 20 tps tg and 200tps pp, while maintaining high scores o
59.
▲
by
tarruda
5mo ago
Looking forward to more open weight releases from Qwen, especially 122B and 397B.
60.
▲
by
tarruda
5mo ago
And as long as Bun doesn't break Claude code, which only uses a subset of it's APIs, this might just pay out.
More ›