Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
liuliu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
liuliu
4mo ago
It solves part of the download issue if they actually delivers a 1-bit whole package (currently their download is around 3.5GiB, still not ideal since FLUX.2 [klein] 4B you can get a package including text encoder ~6 GiB). For speed, no. Dr
32.
▲
by
liuliu
5mo ago
One thing people seems not to acknowledge, and this post made it super clear is that NVIDIA kept their lead extremely well in a few years of very high growth. The TFLOPs, the bandwidth, the interconnect mentioned in this post continues to g
33.
▲
by
liuliu
5mo ago
Probably not really. For gaming, I think probably just need to have a better way to explain visual and what the problem is (collision not done correctly, ways to feedback to LLM's experimentation loop how that should be checked and why
34.
▲
by
liuliu
5mo ago
Only if you think B is an important thing. He is easily > $100M from Tesla.
35.
▲
by
liuliu
5mo ago
Since the frontier is only 8-month ahead of DeepSeek, it is hard to see how model training can be a moat as all the tricks are available from open labs in China. You really just need <100m to bootstrap at this point.
36.
▲
by
liuliu
5mo ago
DSv4 generates much faster on NVIDIA class hardware. It is just a very efficient model.
37.
▲
by
liuliu
5mo ago
I am not sure where this comment is from (possibly without looking at this project?). This project is running quasi-frontier model at reasonable tps (~30) with reasonable prefill performance (~500tps) with a high-end laptop. People simply p
38.
▲
by
liuliu
5mo ago
Thanks. I think it is a good explanation, but also suggests a gap. QAT to me, if done right, is the only way to recover performance for extreme quantization regime. The only thing matters of course, if whether it can work. My confidence in
39.
▲
by
liuliu
5mo ago
I am actually getting interested in QAT these days, especially for LSQ+ type, but it doesn't seem like people have done that enough in open-source world at least, for 2-bit / 3-bit OPD with LSQ+ basically.
40.
▲
by
liuliu
5mo ago
The competition is on DeepSeek v4 Flash for similar size / deployment target.
41.
▲
by
liuliu
5mo ago
Agreed. It is nonsensical to argue that a 3B transformer that hard-capped to decode 100 tokens is "intelligent". Of course when we are evaluating whether "transformers" is intelligent or not, we are talking about taking
42.
▲
by
liuliu
5mo ago
> but transformers are not AGI, and they will never be AGI Like the claim "transformers are AGI", this needs proof, otherwise should be prefixed "I think". And honestly, positive proof is easier than negative proof (y
43.
▲
by
liuliu
6mo ago
I think the parent comment is specifically addressing why the black box (or stochastic?) optimizer he used not working.
44.
▲
by
liuliu
6mo ago
Like you said, yes, it is about courage. I just felt that I won't have that courage when I were in his shoes. We can just be different.
45.
▲
by
liuliu
6mo ago
That's a lawful FBI. This is a lawless executive branch. As we all know by now, executive branch has a lot a power that cannot be limited by Congress nor the Courts and erasing a few zeros from 4T market valuation is a piece of cake (a
46.
▲
by
liuliu
6mo ago
I honestly don't know. tim@apple.com is unavailable for quite some time now (since I tried a few years ago), while lisasu@amd.com still works around that time frame.
47.
▲
by
liuliu
6mo ago
He also donated to Kamala Harris campaign. He would also donate to the next Democratic president for their inauguration if they still choose to do this corruptive thing. And your point is?
48.
▲
by
liuliu
6mo ago
Trump is the president. People voted him into the Office. Tim Cook didn't give him the golden statue before he is in the Office. Everyone in the United States is complicit to the horrible things done by the Trump administration by your
49.
▲
by
liuliu
6mo ago
I32 are 8 4-bit value packed into one int32.
50.
▲
Making Apple Neural Engine work in a custom inference stack
(engineering.drawthings.ai)
1 points
by
liuliu
6mo ago
|
0 comments
51.
▲
by
liuliu
6mo ago
Hold my beer: https://imgur.com/a/sNAoghL
52.
▲
by
liuliu
6mo ago
Of course these numbers are ridiculous. Mac Mini (let's assume Apple releases M5 Pro) tops Int8 (let's assume it is the same as FP8, which it is not) at ~50 TFLOPs, with Draw Things, we recently developed hybrid NAX + ANE inferenc
53.
▲
by
liuliu
6mo ago
In case someone don't know, this is the full text: > 2.5.2 Apps should be self-contained in their bundles, and may not read or write data outside the designated container area, nor may they download, install, or execute code which i
54.
▲
by
liuliu
6mo ago
ANE is OK, but it pretty much needs to pack your single vector into at least 128. (Draw Things recently shipped ANE support inside our custom inference stack, without any private APIs). For token generation, that is not ideal, unless you ar
55.
▲
by
liuliu
6mo ago
That is informative, thanks! Yes, I observe the same thing as the model tends to give up (like you said, "dump TMA for a slower fallback") and needs active steering to get good results. But it indeed works further than one-shot fr
56.
▲
by
liuliu
6mo ago
I am not sure you set it up right. Did you have a runnable WolframLanguage file so it can compare results? Did you give it H100 / H200 access to compile and then iterate? My experience is that once you have these two, it does amazing k
57.
▲
by
liuliu
6mo ago
It is a bit insidious that the price hike coincide with the end of 2x promotion, which makes the usage change a bit more obscure.
58.
▲
by
liuliu
6mo ago
It does feel like also impact the usage meter for subscription plans?
59.
▲
by
liuliu
6mo ago
Brand recognition. Since "model-is-the-service", various previously-interesting companies become thin API resellers and the moat is between "selling a dollar for fifty cents" and Brand awareness. I am not saying this in
60.
▲
Show HN: Metal Quantized Attention on M5 Max
(releases.drawthings.ai)
4 points
by
liuliu
6mo ago
|
0 comments
More ›