Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
syntaxing
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
syntaxing
8d ago
Thanks! I really like how the author packaged everything into a container. Definitely going to give it a go over the weekend!
2.
▲
by
syntaxing
8d ago
Halogen as in this right? https://github.com/peonist-ai/halogen-flash-server
3.
▲
by
syntaxing
8d ago
Thanks! Have you seen issues with quantizing the kv cache?
4.
▲
by
syntaxing
8d ago
Vulkan, I have never used ROCm on it but have been debating since the latest big update. How is your prefill? Do you hit over 1K? If it’s 1000K prefill, and 40 TG, I might have to try this over the weekend. Also, can you fit 128K without of
5.
▲
by
syntaxing
8d ago
Can you point me towards the model you use, both the main model and the flash model? Curious if I can get ~30 with a higher quant.
6.
▲
by
syntaxing
8d ago
I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.
7.
▲
Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
(byteshape.com)
86 points
by
syntaxing
8d ago
|
23 comments
8.
▲
by
syntaxing
8d ago
> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of
9.
▲
by
syntaxing
8d ago
I said this before but I wonder if Dan Kan will reboot Atrium. Rally up some old partners and hope Anthropic buys them out for a couple billion.
10.
▲
by
syntaxing
10d ago
It’s not obvious but you can use this with your own local (or any) models. https://support.mozilla.org/en-US/kb/smart-window-byom
11.
▲
by
syntaxing
12d ago
I would pay a good chunk of money if Apple released a local AI hub to coordinate all AI usage locally (including photo indexing).
12.
▲
by
syntaxing
13d ago
Wow thanks for the link. I have zoom on my personal laptop which isnt ideal. I always wanted to run it sandboxed
13.
▲
by
syntaxing
13d ago
Has anyone have good success using AI generated CAD parts? I’ve been trying but it’s always 95% there, but with all hardware, you need 100% right. It’s often quicker and cheaper for me to do it by hand (but I was a mechanical design enginee
14.
▲
by
syntaxing
16d ago
Surprised no one is talking about it but the 0.1 version bumped the parameters from 284B to 552B but “more efficient”, particularly kv cache usage
15.
▲
by
syntaxing
17d ago
I wonder if that’s why 3.8 got so much better? Mixing the reasoning traces from both sides seems to be effective.
16.
▲
by
syntaxing
17d ago
I’m surprised they allow open lid drinks in the lab. One wrong bump and poof 300K easy.
17.
▲
by
syntaxing
17d ago
> 23 agents total. This hit a bit too close to home. Sol has the same issue, spawns a lot of agents for no good reasons (besides burning tokens).
18.
▲
by
syntaxing
18d ago
I’m more curious how each 4 bit quant compares. It seems like NVFP4 outperforms Q4_K_M in terms of speed and top 1 but is only good for expensive Nvidia cards
19.
▲
by
syntaxing
21d ago
Qwen 3.8 27B is the real deal BUT remember to use froggeric template and/or medium reasoning. https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
20.
▲
by
syntaxing
22d ago
I swear, Qwen 3.8 27B @ Q8 is smarter than Sonnet 5 most of the time. Why wouldn’t corporate America self host at this point, especially with better options like Deepseek Flash and GLM 5.3 flash that’s a middle ground between Sonnet and Opu
21.
▲
by
syntaxing
26d ago
I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
22.
▲
by
syntaxing
26d ago
How do people bypass captcha or robot checks? All I wanted is a price aggregator but it always gets blocked by major retailers.
23.
▲
by
syntaxing
1mo ago
If you get a children’s card, they print your kids name right on it. It’s a nice little souvenir to keep.
24.
▲
by
syntaxing
1mo ago
Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored t
25.
▲
by
syntaxing
1mo ago
I’m more curious on the size. If it’s smaller than or equal size to GLM 5.3, this would be a crazy good model. If it’s closer to deepseek pro, it would be a good model. If it’s near Kimi K3, I think it’s competitive but nothing particularly
26.
▲
by
syntaxing
1mo ago
With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
27.
▲
by
syntaxing
1mo ago
Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.
28.
▲
by
syntaxing
1mo ago
Agreed, and you can write cool infra agents to do stuff for you that runs during off hours like nightly tests and triage.
29.
▲
by
syntaxing
1mo ago
GLM series has made it very practical to self host. If the new update for Deepseek flash holds up, I think it would be silly for some companies to not self host.
30.
▲
by
syntaxing
1mo ago
It’s wild how broken search has become, LLM made it worse but it was already downhill prior to that. Even Reddit forces you login to search within a subreddit and displays the annoying “Best place on internet” overlay after staying on the s
More ›