Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
syntaxing
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
121.
▲
by
syntaxing
5mo ago
Q8 or Q6_UD with no KV cache quantization. I swear it matters even more with small activated parameters MOE model despite the minimal KL divergence drop
122.
▲
by
syntaxing
5mo ago
Llama cpp is vastly superior. There was this huge bug that prevented me from using a model in ollama and it took them four months for a “vendor sync” (what they call it) which was just updating ggml which is the underpinning library used by
123.
▲
by
syntaxing
5mo ago
1. Qwen is mostly coding related through Opencode. I have been thinking about using pi agent and see if that works better for general use case. The usefulness of *claw has been limited for me. Gemma is through the chat interface with lmstud
124.
▲
by
syntaxing
5mo ago
Ironically, even though I write C/++ for a living, I don’t use it for personal projects so I can’t say how well it works for low level coding. Python works great but there’s a limit on context size (I just don’t have enough RAM, and I
125.
▲
by
syntaxing
5mo ago
Inferencing is straight up hard. I’m not accusing them of anything. There’s a crap ton of variables that can go into running a local model. No one runs them at native FP8/FP16 because we cannot afford to. Sometimes llama cpp implementa
126.
▲
by
syntaxing
5mo ago
1. What do you mean by accuracy? Like the facts and information? If so, I use a Wikipedia/kiwx MCP server. Or do you mean tool call accuracy? 2. 3.6 is noticeably better than 3.5 for agentic uses (I have yet to use the dense model). Th
127.
▲
by
syntaxing
5mo ago
Been using Qwen 3.6 35B and Gemma 4 26B on my M4 MBP, and while it’s no Opus, it does 95% of what I need which is already crazy since everything runs fully local.
128.
▲
by
syntaxing
5mo ago
Yes and no. Are you using open router or local? Are the models are good as Opus? No. But 99% of the time, local models are terrible because of user errors. Especially true for MoE, even though the perplexity only drops minimal for Q4 and q4
129.
▲
by
syntaxing
6mo ago
I don’t doubt it is. End of the day, it’s a fine tuned Kimi. They tried to hide it and making their work sound more impressive than it is. It’s easy to have stuff be cheap when you don’t have to train your own model from scratch.
130.
▲
by
syntaxing
6mo ago
What does heavy RL even mean…similar to how the CEO of cursor said how much better the perplexity got when it’s a terrible metric for model fine tune performance? Let’s be real here, it’s Kimi 2.5 fine tuned for Cursor. There’s nothing wron
131.
▲
by
syntaxing
6mo ago
60B for Composer 2…that is built from Kimi K2… what ever happened to “Grok being the best”?
132.
▲
by
syntaxing
6mo ago
With GitHub and Anthropic reducing subscription features, Chinese providers are looking more and more tempting.
133.
▲
by
syntaxing
6mo ago
As the engineering saying goes, nothing more permanent than a temporary solution
134.
▲
by
syntaxing
6mo ago
Curious how it works in other countries, do employees get a portion of the payout?
135.
▲
by
syntaxing
6mo ago
3.4B in 4.5 months…is that all going to Anthropic? Makes it seem so with the wording and how they’re pivoting to Codex too
136.
▲
by
syntaxing
6mo ago
> And the men that had spent longer looking after babies showed the largest drops in testosterone. Those that shared a bed with their infants also had lower levels. Dad here. Maybe…it’s the lack of sleep? Involved fathers tend to have le
137.
▲
by
syntaxing
6mo ago
I went to college as a MechE so unsure if compsci was different. But overall, all the “fun” projects were labs. We have three semesters of hell and all 3 semesters had 2-3 labs, and we write 20 pages or so for EACH lab a week (usually a tea
138.
▲
by
syntaxing
6mo ago
Is it worth running speculative decoding on small active models like this? Or does MTP make speculative decoding unnecessary?
139.
▲
by
syntaxing
6mo ago
I run lmstudio now and it’s more like a “chat” bot. Where as Gemini app is more like an agent.
140.
▲
by
syntaxing
6mo ago
Any way to run this on Gemma 4 only? If there was a “local” mode, I would seriously think about installing this.
141.
▲
by
syntaxing
6mo ago
Which is actually very impressive. 5X is a good deal, given how much more R&D and economy of scale goes into lithium batteries. Flow batteries have 2-3X more cycles and way safer
142.
▲
by
syntaxing
6mo ago
Kinda crazy, it really felt like Meta had the lead in LLMs, especially during the early LLaMa days. What happened for them to fall so far behind? I don’t get how LLaMa 4 was such a big train wreck and they couldn’t correct the course like G
143.
▲
by
syntaxing
6mo ago
I’m genuinely surprised. I use copilot at work which is capped at 128K regardless of model and it’s a monorepo. Admittedly I know our code base really well so I can point towards different things quickly directly but I don’t think I ever ne
144.
▲
by
syntaxing
6mo ago
That’s why I’m a huge proponent of pushing the idea of engineering discipline. And like any organization, discipline comes from the top to bottom. Coming from bottom to top is just a recipe for disaster and clear tell sign of misaligned obj
145.
▲
by
syntaxing
6mo ago
Fun read but a bit too much fluff? I was a design MechE for about a decade and I’m by no means as successful as the OP. But I have worked on a 5 person design team that had an annual net revenue of 30M and another job where I worked on 25M+
146.
▲
by
syntaxing
6mo ago
From what I understand, only works with Tinygrad. Which is better than nothing but CUDA or Vulkan on pytorch isn’t going to work from this. [1] https://docs.tinygrad.org/tinygpu/
147.
▲
by
syntaxing
6mo ago
Wasn’t Composer 2 a “fine tune” of Kimi2.5?
148.
▲
by
syntaxing
6mo ago
Are the dedicated GPU cards on another machine or you’re using eGPU with the framework?
149.
▲
by
syntaxing
6mo ago
Have you used it with any agents or claw? If so, which model do you run?
150.
▲
by
syntaxing
6mo ago
Wow this is super interesting. This creates a local “Gemini” front end and all. This is more or less a generative AI aggregator where it installs multiple services for different gen modes. I’m excited to try this out on my strix halo. The b
More ›