Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kamranjon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
121.
▲
by
kamranjon
3mo ago
Antirez has a GLM 5.2 branch now in dwarfstar: https://github.com/antirez/ds4/tree/glm5.2 It heavily utilizes ssd streaming from my understanding and I think he mentioned getting some semi usable speeds on a
122.
▲
by
kamranjon
3mo ago
Not sure it’s a blog post really but I saw a video of Rutger Bregman talking about moral ambition and it directly led to me applying at a non-profit that was better aligned with my own ideals and leaving my job at a pretty large and morally
123.
▲
by
kamranjon
3mo ago
Wanted to share Antfly which I think serves a similar niche: https://antfly.io/ https://github.com/antflydb/antfly They’ve put a lot of effort into optimizing the local llm pipelines and I have a lot o
124.
▲
by
kamranjon
3mo ago
It’d be real funny if this was just GLM 5.2 trained on Cursor data
125.
▲
Tencent Releases Hy3 295B parameter open weight model
(huggingface.co)
3 points
by
kamranjon
3mo ago
|
0 comments
126.
▲
by
kamranjon
3mo ago
This would be a pretty cool addition to the duckdb HNSW search project I found on here some time ago: https://github.com/jasonjmcghee/portable-hnsw What I think is really cool is that the search happens using http rang
127.
▲
by
kamranjon
3mo ago
One of the big benefits I’ve heard about for the framework, especially with the Noctua cooler, is that it’s virtually silent.
128.
▲
by
kamranjon
3mo ago
In case it saves anyone some time (from the article): "The AMD Ryzen AI Max+ 395(Strix Halo) processor has been available since Spring 2025 and the Halo doesn’t offer anything new on that front." It has the same 256 GB/s memo
129.
▲
by
kamranjon
3mo ago
thank you for the tip here! would you say tinc can work more or less like tailscale? I saw this: "As long as one node in the VPN allows incoming connections on a public IP address (even if it is a dynamic IP address), tinc will be able
130.
▲
by
kamranjon
3mo ago
How common is it for peer reviewed papers like this to be so far off their claimed findings? “According to Google's peer-reviewed and published paper, they claim to have a true positive rate (TPR) above 99.97% -- meaning that they will
131.
▲
by
kamranjon
3mo ago
This is very cool - I will likely see if I can use it in place of tailscale for my local LLM hosting. I feel like not having that required login would be great. Also the direct connect feature seems pretty cool, since that’s usually all I n
132.
▲
by
kamranjon
3mo ago
Curious what data sets you used?
133.
▲
by
kamranjon
3mo ago
It takes longer to drive the length of Sweden than it does the length of California.
134.
▲
by
kamranjon
3mo ago
I'm really happy this is one of the top comments here, I am fully local as well. Just wanted to leave a note for folks who might not have the memory to run a big 32gb model - I just found out there are some pruned models that have real
135.
▲
by
kamranjon
3mo ago
ah yeah you're correct - sorry for the confusion
136.
▲
by
kamranjon
3mo ago
This is interesting, I haven’t actually heard you suggest that the labs are focusing on this benchmark before. Have you come around to this position as a result of the quality of pelicans you’ve been getting? The reason I thought this was a
137.
▲
by
kamranjon
3mo ago
This is the way
138.
▲
by
kamranjon
3mo ago
I completely disagree, it is probably the best platform currently for this - and the way I run it is as a server with tailscale accessible from my coding machine (same as you suggest here) - the difference is that you can stop the server, u
139.
▲
by
kamranjon
3mo ago
If anyone is interested in doing something seriously useful with these neural cores, there is this incredible write up on getting ModernBERT running on them: https://stephenpanaro.com/blog/modernbert-on-apple-neural-en.
140.
▲
by
kamranjon
3mo ago
I think the 9b and 31b dense are Gemma models and the 35B-MoE, and 397B-MoE are Qwen models since these are model sizes covered by each of them respectively
141.
▲
by
kamranjon
3mo ago
Strix halo I believe is 256GB/s max memory bandwidth and M5 Max is 614GB/s - M3 Ultra is up to 800gb/s
142.
▲
by
kamranjon
3mo ago
Ah yea after watching one of the creators youtube videos I realize these benchmarks are combining prefill and decode which isn't super helpful - it seems this struggles with the exact same bottlenecks as all strix halo setups, memory b
143.
▲
by
kamranjon
3mo ago
I generally agree for everything except Macbook Pros which outperform most available desktop setups for AI tasks - but they are also now out of reach for most people after the price hikes (6.7k now for 128gb, i got mine for 4.7k just about
144.
▲
by
kamranjon
3mo ago
A fine-tuned classifier purpose fit for a specific task can easily outperform a SOTA LLM on more modest hardware and often makes a lot more sense.
145.
▲
by
kamranjon
3mo ago
Benchmarks are here: https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/ Would love to see DeepSeek V4 flash/pro and MiniMax M3 benchmarks but already these are pretty impressive, first strix Halo setup I'
146.
▲
by
kamranjon
3mo ago
This was the easiest link I could find but TechCrunch also has an article on the departure - this one didn’t require turning off ad blockers though and has a bit more rumor mill stuff: https://chinabizinsider.com/alibabas-qw
147.
▲
by
kamranjon
4mo ago
Thanks for the clarification - Google does publish more than others - and I actually really appreciate the work they are doing with the Gemma models, which are truly competitive open models. I do wish they’d publish more in depth papers on
148.
▲
by
kamranjon
4mo ago
Google hasn’t published much in depth ML work since T5 (which was hugely influential at the time) - most Gemma releases are 1-3 page model card pdfs these days with no in depth analysis. Even TurboQuant is shaking out to have basically been
149.
▲
by
kamranjon
4mo ago
The paper actually references testing their DSpark speculative decoding strategy with Qwen 3 4b, 8b and 14b models so while I doubt they will release builds themselves, they’ve open sourced (DeepSpec) their training pipeline for this so we
150.
▲
by
kamranjon
4mo ago
There was a recent exodus from Qwen of researchers who supported their open source efforts, I’m not sure we will see many new open models from them past the 3.6 series.
More ›