Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
antirez
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
61.
▲
by
antirez
5mo ago
I agree with you, practically. But there is another angle of the story: for instance models are starting to be useless to do security stuff, since they are every day more censored. Also prices skyrocketed in the latest months, what will hap
62.
▲
by
antirez
5mo ago
Mmmm, nope if you do the smart thing. MacBook M5 max 128gb is a premium laptop at 6k, but with it you can do many things and is your good main driver for the day. Then, it can also run DeepSeek V4 flash and perform non trivial tasks locally
63.
▲
by
antirez
5mo ago
Didn't know M2.7 could also resist extreme quantizations, I had the feeling that being it shipped Q8 it was easily damaged in that way. Very interesting data point! And thank you for the nice words. Btw it really looks like ~250/3
64.
▲
by
antirez
5mo ago
Btw, a few data points: 1. DS4F can run on a 128GB MacBook. M2.7 is larger (8 bit weights of routed experts). There is to see how it holds at 4 bits. At 2 bits it may not work well at all. 2. Just the KV cache of M2.7 would take ~50GB for 2
65.
▲
by
antirez
5mo ago
In the same page DS4F scores much better on Omniscent Accuracy. I would take those numbers with a bit of salt. For instance I ran different benchmarks against Qwen 3.6 27B and DS4F quantized at 2bit. DS4F hallucination rate is much lower.
66.
▲
by
antirez
5mo ago
May I ask you what did you used for the DS4F inference? It is a model with very low hallucination rate in my tests.
67.
▲
by
antirez
5mo ago
Nope, with the anti-refusal vector loaded you can ask many things for instance related to computer security and if you want to learn, it is a lot better of a model that continuously says you "I can't help you with this problematic
68.
▲
by
antirez
5mo ago
Not just that. The other day I was able to ask DeepSeek v4 (with the anti-refusal vector loaded) all the top tricks to steal a lollypop to a child.
69.
▲
by
antirez
5mo ago
Send patches! But remember that many speedups end being not exactly correct and the logits drift. But there is extensive testing and even ds4-eval now to test how it performs.
70.
▲
by
antirez
5mo ago
Yep, the code overlap is minimal, a few kernels. Some quantization code for the quantizer it implements. DwarfStar 4 is not a fork of llama.cpp, but without llama.cpp the project would be a lot more lacking, since I was able to get all the
71.
▲
by
antirez
5mo ago
Runs on 96GB MacBooks. 128GB is better. Check the README of DwarfStar.
72.
▲
by
antirez
5mo ago
Thank you for posting this! Just a clarification, with DwarfStar steering features I was able to completely remove refusal from DS4. It is only the example dataset (prompt pairs I provide) which is a toy, not the abilities. I thought that w
73.
▲
by
antirez
5mo ago
Check the readme better. The code overlap with ggml is very small, but a few kernel and ideas and the quants code were taken. Still the project connection with llama.cpp and ggml is huge and also present in the license because it's not
74.
▲
by
antirez
5mo ago
Prefill is 400 t/s in that hardware. Just if the prompt is very short you can't see the real speed and it will default to single token context processing.
75.
▲
by
antirez
5mo ago
The bugs were on the API tool call handing. The model worked well. I would retest with updated code and gguf and I and many others never saw it missing anything obvious. Reliable tool call and reasoning. The project is a few days old so cer
76.
▲
by
antirez
5mo ago
It works on your computer I believe. There are a few positive reports.
77.
▲
by
antirez
5mo ago
You misunderstood the OP. I hinted, in my blog, at my interest to also putting an agent harness inside.
78.
▲
by
antirez
5mo ago
There is no need for another agent, functionally. But if you follow the idea of DS4 itself: the API agents use forces to do odd things, like translating the DSML stanzas to JSON, with all the canonicalization / KV cache checkpointing p
79.
▲
by
antirez
5mo ago
It's a mix of extreme sparsity but with the routed expert doing a non trivial amount of work (and it is q8), and projections and routing not being quantized as well. Also the fact it's a QAT model must have a role I guess, and I q
80.
▲
by
antirez
5mo ago
Yep that happens with coding agents sending a very large system prompt. And also when later tool calling feed it large files or diffs. But with the M3 ultra the prefill speed is almost 500 t/s that is quite into the very usable zone. W
81.
▲
by
antirez
5mo ago
It runs both q2 and original (4 bit routed experts). At the same speed more or less. The q2 quants are not what you could expect: it works extremely well for a few reasons. For the full model you need a Mac with 256GB.
82.
▲
by
antirez
5mo ago
DS4 can process 460 prompt tokens per second. Not stellar but not so slow. On M3 max. See the benchmarks on readme.
83.
▲
by
antirez
5mo ago
True quantitatively, not qualitatively. DeepSeek V4 is not capable of doing what a human brain can do, of course, but for the tasks it can do, it can do it at a speed which is completely impossible for a human, so comparing the two requires
84.
▲
by
antirez
5mo ago
A random, funny, interesting and telling data point: my MacBook M3 Max while DS4 is generating tokens at full speed peaks 50W of energy usage...
85.
▲
by
antirez
5mo ago
It was not my best (nor normal) behavior, but the point in this case is that the OP offered very little in his rebuttal. A more contextualized reply would have improved mine as well. I believe actually the person that published this LLM cou
86.
▲
by
antirez
5mo ago
You are right, sorry the name is very similar and I thought it was: https://x.com/angeloskath
87.
▲
by
antirez
5mo ago
Google the name of the author.
88.
▲
by
antirez
5mo ago
Context: he is one of the MLX developers, a skilled ML researcher.
89.
▲
by
antirez
5mo ago
If Pro is the same model (hard to tell, I'm not sure) it has a token budget to think (test time scaling) which is huge compared to the Codex endpoint.
90.
▲
by
antirez
5mo ago
Redis was completely built in this way since the start. I believe this is a better way to create software. Compromise in design is, in my opinion, something to avoid: feedbacks are important, but often times a single person that studied a l
More ›