Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dnhkng
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
dnhkng
7mo ago
Yes! I tried that pretty early on, the its basically never good. Its described in the the section: https://dnhkng.github.io/posts/rys/#the-beginning-of-llm-neu...
32.
▲
by
dnhkng
7mo ago
Yes, I was using Base64 to 'jailbreak' LLMs back in the day (so similar), and thats what led me to the hypothesis, and months of GPU use to find optimal later dultication!
33.
▲
by
dnhkng
7mo ago
Because its generally expected that models only work 'in distribution', i.e. they work on stuff they have previously seen. They almost certainly have never seen regular conversations in Base64 in their training set, so its weird t
34.
▲
by
dnhkng
7mo ago
Here's an extract, the core TL;DR for a feel of the article. "And now for the weirdness: There was never the case where any Transformer layer would have seen the output from a future layer! Layer 10 is trained on layer 9’s output
35.
▲
by
dnhkng
7mo ago
Thanks!
36.
▲
by
dnhkng
7mo ago
Yes, I've tried duplicating indvidual layers, but its not useful. I think this hasn't been tried before because it's totally unintuitive that feeding the output from later layers into previous ones would actually do anything.
37.
▲
by
dnhkng
7mo ago
Maybe, but the interesting thing for me it this only works with specific 'chunks' of the transformer layer stack. More or less that the optimal leads to worse performance.
38.
▲
by
dnhkng
7mo ago
Have a look at the boundaries in the heatmaps. They are of course open to interpretation, but it suggest to me that the models develop 'organs' for processing different types of data, and without duplicating the 'whole organ&
39.
▲
by
dnhkng
7mo ago
No worries, happy to discuss anyway :) MoE (mixture of experts), is an architecture that forces sparsity (not all 'neurons' are active during the forward pass. This is pretty much orthogonal to that; it works with dense and MoE mo
40.
▲
by
dnhkng
7mo ago
Agrees, but one thing to note: I really think from the experiments that 'organs' (not sure what to term this), develop during massive pretraining. This also means maybe looping the entire models is actually not efficient. Maybe a
41.
▲
by
dnhkng
7mo ago
I build my own analysis tools. I'm just finishing up running the current generation of LLMs (MiniMax M2.5 and the Qwen3.5 family), and then I will put it all on Github. It less 'tool', than an assorted set of scripts, tailore
42.
▲
by
dnhkng
7mo ago
Thanks for the link! I think that these models have to learn to efficiently use their parameters, and the best way to do that is 'evolve' (yes, a bad word for it), structures over pretraining time. Unfortunately, they don't h
43.
▲
by
dnhkng
7mo ago
I did, but the combinatorics are mad. I have also tried training a meta-model that predicts the outputs of the combinations. I will make another post if the topic is popular; its pretty geeky though, even more than my usual blog posts...
44.
▲
Show HN: How I topped the HuggingFace open LLM leaderboard on two gaming GPUs
(dnhkng.github.io)
495 points
by
dnhkng
7mo ago
|
126 comments
45.
▲
by
dnhkng
10mo ago
She even read the article! (she did not find it funny though ;) )
46.
▲
by
dnhkng
10mo ago
Sure, its a free country! (I'm Australian, living in Germany).
47.
▲
by
dnhkng
10mo ago
I've had a bit of practice, but I don't have the right gear for this level of soldering. It took maybe an hour to solder in 2 components, after many failed attempts. Persistence beats intelligence?
48.
▲
by
dnhkng
10mo ago
No, it was done over the course of weeks, and I'm not motivated enough to do the production work required for good quality videos.
49.
▲
by
dnhkng
10mo ago
Running LLM's directly might not be effective. I think there are probably Law Firms/doctors offices that would gladly pay ~3-4K euro a month to have this thing delivered and run truely "on-prem" to work with documents th
50.
▲
by
dnhkng
10mo ago
Fair points, but the deal is still great because of the nuances of the RAM/VRAM. The Blackwells are superior on paper, but there's some "Nvidia Math" involved: When they report performance in press announcements, they do
51.
▲
by
dnhkng
10mo ago
These are on a custom board from Nvidia, so its not possible to separate them. I think the seller usually gets H100's and them into a custom case, with a PCIE adapter to the server GPUs. This thing too unwieldy to make into a desktop (
52.
▲
by
dnhkng
10mo ago
I'm downloading DeepSeek-V3.2-Speciale now at FP8 (reportedly Gold-medal performance in the 2025 International Mathematical Olympiad and International Olympiad in Informatics). It will fit in system RAM, and as its mixture of experts a
53.
▲
by
dnhkng
10mo ago
This is hard to say for sure. I had 4x 4090, that I had bought for about $2200 each in early 2023. I sold 3 of them to help pay for the GH200, and got 2K each.
54.
▲
by
dnhkng
10mo ago
Correct! I added an Nvidia T400 to the rig recently, as it gives me 4x Display ports, and a whole extra 2GB VRAM!
55.
▲
by
dnhkng
10mo ago
It was fun when the seller told me to come and look in the back of his dirty white van, because "the servers are in here". This was before I had seen the workshop etc.
56.
▲
by
dnhkng
10mo ago
Oh no, thats not right. 20 Kg was in the original server case. With the Aluminium frames, and glass panel, its more like 40 Kg now... Shit, maybe I should take it off the Lack table...
57.
▲
by
dnhkng
10mo ago
lol, I tried posting stuff on Twitter, but never got any traction. This might be too nerdy for that crowd?
58.
▲
by
dnhkng
10mo ago
I'll find out soon, but without this hack, the GPUs are non-functional.
59.
▲
by
dnhkng
10mo ago
Makes sense. I'm so used to the naming I forgot it's not common knowledge. I hope the new title is clearer.
60.
▲
I got an Nvidia GH200 server for €7.5k on Reddit and converted it to a desktop
(dnhkng.github.io)
373 points
by
dnhkng
10mo ago
|
108 comments
More ›