Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gdiamos
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
gdiamos
10d ago
I wouldn’t trust an AI company to honor this as far as I could throw them
2.
▲
by
gdiamos
16d ago
How to steal ideas with AI. step 1, identify high value users by net worth, citation count, or number of followers step 2, select all prompts by high value users step 3, invest 10 billion thinking tokens in modeling an objective for each us
3.
▲
by
gdiamos
16d ago
Training improvements are very easy to copy.
4.
▲
by
gdiamos
18d ago
Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models Training a model this small on them is distillation When models of this size were last studied seriously such corpora
5.
▲
by
gdiamos
18d ago
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end
6.
▲
by
gdiamos
18d ago
blog: https://gregdiamos.com/2026/09/07/outrageously-small-neural-... X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20 I added some of the main points to th
7.
▲
Outrageously Small NNs: Emergent Reasoning at 6,616 Tok/s on One Intel AMX Core
(gregdiamos.com)
1 points
by
gdiamos
18d ago
|
0 comments
8.
▲
by
gdiamos
19d ago
I liked how the article was aimed at Mark personally. Not everyone gets super voting shares, but everyone gets a life and has to live on the same planet.
9.
▲
by
gdiamos
19d ago
OxyContin, Enron, WorldCom, Super-size-me, Pets.com, Asbestos in the ceiling tiles... It was always burning since the world's been turning
10.
▲
by
gdiamos
19d ago
we will look back on it as the big tobacco of our generation
11.
▲
by
gdiamos
19d ago
how's the battery life?
12.
▲
Outrageously Small Neural Networks: 6,616 tok/s on One Intel AMX Core [pdf]
(huggingface.co)
2 points
by
gdiamos
19d ago
|
1 comments
13.
▲
by
gdiamos
19d ago
I think we should revisit outrageously small neural nets. I needed a cheap model that runs at over 10k token/sec on a single CPU core for some data processing. So I gave Anthropic claude code a pile of tokens to build one. It made thre
14.
▲
by
gdiamos
19d ago
Progress compared to SLMs and the early days of deep learning is real. However, I know of no theoretical limits on scaling laws other than compute and data.
15.
▲
by
gdiamos
23d ago
I think it means that we should be aiming further ahead
16.
▲
by
gdiamos
26d ago
I’d like to see more of these models. I’ve been using diffusion Gemma and it is very fast on GPUs in output token/sec. In the diffusion Gemma whitepaper, they say they could have done better with more time and compute. Even with those
17.
▲
by
gdiamos
1mo ago
Best case scenario
18.
▲
by
gdiamos
1mo ago
There's certainly a place for enterprise and not breaking what's working. Shouldn't that be 0 innovation tokens though?
19.
▲
by
gdiamos
1mo ago
In hindsight I disagree. Instead I like “only work on impossible problems” Most of them turn out to be impossible, but some of them turn out to be possible. I’ve never met anyone who could pick 3 and be confident in getting even one right.
20.
▲
by
gdiamos
2mo ago
How big is the open model? 30B?
21.
▲
by
gdiamos
2mo ago
Christopher Nolan beat you to it
22.
▲
by
gdiamos
2mo ago
vLLM is originally marketed as paged attention, but in hindsight, separating the web server and GPU process, continuous batching, kv caching / chunking, and a huge model library including low precision mattered more. I wonder how much
23.
▲
by
gdiamos
2mo ago
I wish I could get a model to state its assumptions.
24.
▲
by
gdiamos
2mo ago
How do you ban melted sand?
25.
▲
by
gdiamos
2mo ago
I think we should shut it off. It would force US companies to build open models.
26.
▲
by
gdiamos
2mo ago
I want a hosted paper to be archived. That means that 10 years from now I don’t want think about making sure the hosting server is up. I also want it to have a standard format for bibliography, DOI, and authors. I agree it isn’t much, but i
27.
▲
by
gdiamos
2mo ago
thank god, these parameters are so confusing
28.
▲
by
gdiamos
2mo ago
as soon as you release a way of measuring it, you give LLMs a signal to optimize
29.
▲
by
gdiamos
2mo ago
Being on the review board comes with a promise to not be evil right?
30.
▲
by
gdiamos
3mo ago
No, I want arxiv to host the paper, not to review the paper. I wouldn't want my google drive to start telling me my paper was too sloppy. I just want a link.
More ›