Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
am17an
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
am17an
8mo ago
Honestly you can run this on a 16GB VRAM GPU with llama.cpp. Just try it!
32.
▲
by
am17an
8mo ago
One often overlooked after that is ggml, the tensor library that runs llama.cpp is not based on pytorch, rather just plain cpp. In a world where pytorch dominates, it shows that alternatives are possible and are worthy to be pursued.
33.
▲
by
am17an
8mo ago
Holy smokes we're cooked.
34.
▲
by
am17an
8mo ago
Maintainers time is a more scarce resource than free tokens. I would much rather get my time back after reading those PRs
35.
▲
by
am17an
9mo ago
1) Python is unreadable." Would you prefer C or C++? > Unironically, yes. Unless I never plan to look at that code again
36.
▲
Every LLM hallucinates that std:vector deletes elements in LIFO order
(am17an.bearblog.dev)
6 points
by
am17an
9mo ago
|
1 comments
37.
▲
by
am17an
10mo ago
Use llama.cpp? I get 250 toks/sec on gpt-oss using a 4090, not sure about the mac speeds
38.
▲
by
am17an
10mo ago
Well a 1000 line PR is still not welcome. It puts too much of a burden on the maintainers. Small PRs are the way to go, tests are great too. If you have to submit a big PR, get buy in from a maintainer first that they will review your code.
39.
▲
by
am17an
10mo ago
I agree, this would be in the same vein as "STL returns a verbose type, it's okay to use auto here because no-one cares"
40.
▲
by
am17an
10mo ago
This has got to be a hell of a story if it's true.
41.
▲
by
am17an
10mo ago
Usually codebases disallow auto because without an IDE it's difficult to see the type. I think this reduces the cognitive load of C++ a bit. The only time it is allowed is getting iterators types from STL containers. I remember frettin
42.
▲
by
am17an
10mo ago
All a LLM does is hallucinate, some hallucinations are useful. -someone on the internet
43.
▲
by
am17an
10mo ago
I hope this is not a serious argument. This deal is OOM larger than other deals
44.
▲
by
am17an
10mo ago
It's bad for consumers period. A deal that hampers 40% of global supply shouldn't be a thing, it's predatory. I know DRAM is not a necessity, but considering that PCs are going to be affected means this affects real things li
45.
▲
by
am17an
11mo ago
This is one of my favourite problems, I still remember that it has a very real edge case even though I solved it more than 10 years ago. Thank you for the problem!
46.
▲
by
am17an
11mo ago
The non-thinking version is the best writer by far. Excited for this one! They really cooked some different from other frontier labs.
47.
▲
by
am17an
1y ago
What is parallel decoding?
48.
▲
by
am17an
1y ago
Humbert is a bit on the nose, for those who get the Lolita reference.
49.
▲
by
am17an
1y ago
The answer is 0
50.
▲
A gentle introduction to GEMM using MMA tensor cores
(am17an.bearblog.dev)
4 points
by
am17an
1y ago
|
0 comments
51.
▲
by
am17an
1y ago
I don't feel weird about it because it's a trillion dollar business, and it's in their best interests to muddy the waters regarding any cohesive argument. But when you just look at these things from first principles, boasting
52.
▲
by
am17an
1y ago
The kind of person who doesn’t believe this is bad for children, is the same kind of person who would believe Big Tobacco sponsored studies about how cigarettes don’t cause cancer in the 80s. For the data driven like yourself, I remember a
53.
▲
by
am17an
1y ago
Probably for the better, its incredibly damaging to young brains.
54.
▲
by
am17an
1y ago
This model is literally amazing. Everyone should try to get their hands on a H100 and just call it a day.
55.
▲
by
am17an
1y ago
They must really be having a bad time if Anthropic of all labs is willing to share their infra details. On the actual precision bug, it is quite unfortunate on FMA side, numerical issues are often deeply bewildering and no AI can solve them
56.
▲
by
am17an
1y ago
Somehow that's still an understatement
57.
▲
by
am17an
1y ago
Isn't this factually wrong? Grok-4 used as much compute on RL as they did on pre-training. I'm sure GPT-5 was the same (or even more)
58.
▲
by
am17an
1y ago
The point is those kernels exist already, you can just use them off the shelf. In the case where you're trying to write a production grade kernel without operating at that part of the stack... well good luck with that.
59.
▲
by
am17an
1y ago
There’s a GitHub link which is open from last year, about the missing license in ollama. They have not bothered to reply, which goes to show how much they care. Also it’s a YC company, I see more and more morally bankrupt companies making t
60.
▲
by
am17an
1y ago
This is nice and useful because the new GPT-OSS model uses this technique. Kudos to the original authors!
More ›