Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zargon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
61.
▲
by
zargon
5mo ago
Also from phoronix, a comparison with AMD R9700 and RTX 6000 Ada (because Nvidia has not sent them a blackwell card): https://www.phoronix.com/review/intel-arc-pro-b70/2
62.
▲
by
zargon
6mo ago
I call it autocorrupt :)
63.
▲
by
zargon
6mo ago
It's only in preview right now. And anyway, yes, models regularly get updated training. But in this case, it's more likely just to be a tooling issue.
64.
▲
by
zargon
6mo ago
I think you mean ollama vs llama.cpp.
65.
▲
by
zargon
6mo ago
There is no BF16. There is no FP8 for the instruct model. The instruct model at full precision is 160 GB (mixed FP4 and FP8). The base model at full precision is 284 GB (FP8). Almost everyone is going to use instruct. But I do love to see b
66.
▲
by
zargon
6mo ago
Flash is less than 160 GB. No need to quantize to fit in 2x 96 GB. Not sure how much context fits in 30 GB, but it should be a good amount.
67.
▲
by
zargon
6mo ago
> ~100GB at 16 bit or ~50GB at 8bit quantized. V4 is natively mixed FP4 and FP8, so significantly less than that. 50 GB max unquantized.
68.
▲
by
zargon
6mo ago
That article is a total hallucination. "671B total / 37B active" "Full precision (BF16)" And they claim they ran this non-existent model on vLLM and SGLang over a month and a half ago. It's clickbait keyword sl
69.
▲
by
zargon
6mo ago
The Flash version is 284B A13B in mixed FP8 / FP4 and the full native precision weights total approximately 154 GB. KV cache is said to take 10% as much space as V3. This looks very accessible for people running "large" local
70.
▲
by
zargon
6mo ago
Shall I implement it? no https://gist.github.com/bretonium/291f4388e2de89a43b25c135b4...
71.
▲
by
zargon
6mo ago
For FIM, there's Qwen3 Coder Next. Although Mistral's model card seems to indicate that Devstral 2 doesn't support FIM, it seems very odd that it wouldn't. I have been meaning to test it.
72.
▲
by
zargon
6mo ago
Excellent job with this! I tried a few combinations that completely fail on other calculators and yours gets VRAM usage pretty much spot on, and even the performance estimate is in the ballpark to what I see with mixed VRAM / RAM workl
73.
▲
by
zargon
6mo ago
I just loaded up Qwen3.6 27B at Q8_0 quantization in llama.cpp, with 131072 context and Q8 kv cache: build/bin/llama-server \ -m ~/models/llm/qwen3.6-27b/qwen3.6-27B-q8_0.gguf \ --no-mmap \ --n-
74.
▲
by
zargon
6mo ago
I just tested it and have to make a correction. With llama.cpp, 262144 tokens context (Q8 cache) used 8.7 GB memory with Qwen3.6 27B. Still very impressive.
75.
▲
by
zargon
6mo ago
These calculators are almost entirely useless. They don't understand specific model architectures. Even the ones that try to support only specific models (like the apxml one) get it very wrong a lot of the time. For example, the one yo
76.
▲
by
zargon
6mo ago
Yes, definitely it's the bottleneck for most use cases besides "chatting". It's the reason I have never bought a Mac for LLM purposes. It's frustrating when trying to find benchmarks because almost everyone gives de
77.
▲
by
zargon
6mo ago
Qwen3.5 series is a little bit of an exception to the general rule here. It is incredibly kv cache size efficient. I think the max context (262k) fits in 3GB at q8 iirc. I prefer to keep the cache at full precision though.
78.
▲
by
zargon
6mo ago
Everything is benchmaxxed. Whack-a-mole training is at least as representative of what is getting added to models as more general training advances.
79.
▲
by
zargon
6mo ago
LLMs need diverse and extensive training data to be good at a specific thing. We don't (yet?) know how to train a small model that is really good at one programming language. Just big models that are good at a variety of languages (plu
80.
▲
by
zargon
6mo ago
If someone doesn't specifically say prefill then they always mean decode speed. I have never seen an exception. Most people just ignore prefill.
81.
▲
by
zargon
6mo ago
And I don't drink coffee, just water. Since software is priced as beverage equivalence, logically that means I should get software for free.
82.
▲
by
zargon
6mo ago
Why do you merge the GGUFs? The 50 GB files are more manageable (IMO) and you can verify checksums as you say.
83.
▲
Intel Arc Pro B70 Open-Source Linux Performance Against AMD Radeon AI Pro R9700
(phoronix.com)
5 points
by
zargon
6mo ago
|
1 comments
84.
▲
by
zargon
6mo ago
Previous discussion: https://news.ycombinator.com/item?id=47546732
85.
▲
by
zargon
6mo ago
Click on the timestamp link to go to the comment's own page where it will be rendered black instead of gray.
86.
▲
by
zargon
6mo ago
Comments are marked dead by automatic processes, not through downvotes. They're dead before anyone sees them, and you can't vote on a dead comment. amangsingh's comments have probably triggered some automated moderation. Prob
87.
▲
by
zargon
6mo ago
Swappa is great for sellers and bad for buyers. They use a Paypal marketplace product and have no actual authority in a transaction, they just match buyers and sellers. If the phone is not as described, you'll have to pay return shippi
88.
▲
by
zargon
6mo ago
Yeah. The response to the issue of the LLM cheating should be removing the LLM's access to the ledger. If the architecture allowed the LLM access to the ledger, I have zero reason to believe any amount of cryptography will prevent it.
89.
▲
by
zargon
7mo ago
*Only gamers know that joke.
90.
▲
by
zargon
7mo ago
Yes and yes. NVMe storage is very slow, so it can get away with such things.
More ›