Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lostmsu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
lostmsu
5d ago
It doesn't show any indications of solving catastrophic forgetting.
2.
▲
by
lostmsu
5d ago
This is slop. 8M parameter dense model with context length 64 that you train on enwik9 in 2h will have 1.15 bpb. This model has 1.8 (bits per byte, lower is better).
3.
▲
by
lostmsu
7d ago
According to https://goodstartlabs.com/research/verification-is-the-bottl... it is only 2.6x cheaper than DeepSeek V4.1 Flash, and they did not test v4.0 Flash, which would have been same price.
4.
▲
by
lostmsu
7d ago
> Monarch Hadamard MLP: replaces the dense FFN with three learnable Walsh-Hadamard-initialized Kronecker (Monarch) factor pairs interleaved with per-channel diagonal scales, fixed permutations, a SiLU nonlinearity, and a rank-8 input-con
5.
▲
by
lostmsu
7d ago
MS absolutely has a couple of stock-based incentives.
6.
▲
by
lostmsu
7d ago
> And there are bugs in any Linux distribution that are way too complex to fix even if theoretically possible. So it doesn’t make any difference. It does now with LLMs
7.
▲
by
lostmsu
8d ago
Thank God there's the Internet Archive then. https://archive.org/donate I do.
8.
▲
by
lostmsu
8d ago
Good luck with that.
9.
▲
by
lostmsu
9d ago
That's posttraining. Pretraining is the expensive part.
10.
▲
by
lostmsu
9d ago
> just a list of 64 numbers > remember even a few positions? Sure they could A rough estimate of number of positions across all X move games is X^10. For 15 moves it is hopeless to remember even a relatively small part of them. Typica
11.
▲
by
lostmsu
9d ago
> It's just political theater Do you think the referendum in question is not a political theater? I don't know much, but it stands to me that the EU association is magnitudes more serious.
12.
▲
by
lostmsu
10d ago
> prediction which is closer to memorization > don't memorize inputs - they predict them I feel some tension here. > rice grains on a chess board? Sure, but this has nothing to do with chess, and nothing to do with how many ga
13.
▲
by
lostmsu
10d ago
They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board? The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intellige
14.
▲
by
lostmsu
10d ago
I am surprised nobody brought up Accelerando yet.
15.
▲
by
lostmsu
10d ago
Papers charged per "user" since times immemorial.
16.
▲
by
lostmsu
10d ago
You are saying "No it is not" without an argument. The fact that computer systems could play chess yet not being AGI has no relevance to LLMs' ability to play chess being AGI, because the point is about G, not I. There's
17.
▲
by
lostmsu
10d ago
Do they?
18.
▲
by
lostmsu
10d ago
Yes, a good analogy. Except the cat actually follows the football rules and can beat some humans. And has no physical limitations to play other kinds of sport that you might imply.
19.
▲
by
lostmsu
10d ago
The fact that LLMs can play chess at any level is a strong indication we are in AGI.
20.
▲
by
lostmsu
10d ago
Why?
21.
▲
by
lostmsu
11d ago
So I am building a voice assistant to control AI harnesses, and recently tried switching from GLM 5.3 Flash to Gemini 3.8 Flash because of higher tok/s and better rate limits. Before that I also used Kimi K3 and DeepSeek-V4-Flash-0731.
22.
▲
by
lostmsu
17d ago
Perhaps the commenters don't care. Take solace in that you do.
23.
▲
by
lostmsu
19d ago
This lacks comparison to regular economy.
24.
▲
by
lostmsu
21d ago
You don't rinse them after washing?
25.
▲
by
lostmsu
21d ago
You are as wrong as your parent comment. Not only Russia is invading US, but US is less of the supporter of Russia than EU. EU literally funds Russia and doesn't plan stopping completely.
26.
▲
by
lostmsu
22d ago
It is not just a civil offence. It is potentially an organized crime.
27.
▲
by
lostmsu
22d ago
Yes. Their architecture recomputes every time so at 150k context every request will have to spend 1.5 min waiting for the model to reread the context. Say avg model response length is 1024 tok. At 50 tok/s normal providers do your turn
28.
▲
by
lostmsu
23d ago
5.6.1
29.
▲
by
lostmsu
23d ago
Pure marketing.
30.
▲
by
lostmsu
23d ago
They don't have cache (e.g. KV cache). But they write down what you sent earlier to say they cached it! To still bill the same as uncached later (because they didn't actually cache it)!
More ›