Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nearbuy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
nearbuy
15d ago
They did not hardcode that, because: a) There's zero evidence of them doing so b) Some models released before the car wash problem was discovered would consistently get it right c) Hardcoding it is pointless. No one is seriously asking
2.
▲
by
nearbuy
17d ago
JimDabell said they didn't come up with that thought themselves. They never implied the poster didn't mean it. sm-silversight responded, "How do you know it's not his/her thought?" You can reasonably interpret
3.
▲
by
nearbuy
17d ago
In this comment section alone, 5 people have used some variant of that aphorism, and it's used constantly on hn and other sites. They've heard it.
4.
▲
by
nearbuy
21d ago
> Things like "You're right, but for a stronger reason than you said:" I also get this a lot, but almost always with " and for a stronger reason than you said", which makes it sound more like it's reinforci
5.
▲
by
nearbuy
23d ago
Buckmaster (the mathematician) and Alpöge used and credit AI substantially for their proof. Even if OpenAI did copy their ideas, it still wouldn't show that this didn't come from AI improving. OpenAI's proof is substantially
6.
▲
by
nearbuy
26d ago
You're conflating two very different meanings of "the speed of light". Confusingly, "speed of light" can refer to the universal constant, c, which does not change in glass, or to the speed that light travels in a me
7.
▲
by
nearbuy
29d ago
Why? 1. I didn't say LLMs have made any breakthroughs in math, not because they haven't, but because it's irrelevant to my point. The parent comment is using the same argument academic research opponents have long used agains
8.
▲
by
nearbuy
29d ago
There are several hundred thousand mathematicians producing hundreds of thousands of new results in math each year. Why can't most people name any human contributions to mathematics from the past decade? What is the finished result tha
9.
▲
by
nearbuy
1mo ago
The Kevin Buzzard post linked at the top says they budgeted £1M over 5 years for a smaller proof.
10.
▲
by
nearbuy
1mo ago
I think you're misunderstanding. Astra is at the top of the official ARC-AGI leaderboard, with an ARC-AGI approved harness. It's not a harness specialized for ARC-AGI. It just does the same thing the regular ChatGPT interface does
11.
▲
by
nearbuy
1mo ago
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-
12.
▲
by
nearbuy
1mo ago
I don't think we should count the lower tier models if we're discussing what the top ones are capable of. No one was suggesting that Sonnet is AGI.
13.
▲
by
nearbuy
1mo ago
The last version to fail on those questions was GPT 4.5. Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of
14.
▲
by
nearbuy
2mo ago
It also looks like they're saturating the test, with one LLM hitting the maximum possible score. ( https://www.trackingai.org/home ) The test wasn't made to accurately measure IQs that high.
15.
▲
by
nearbuy
2mo ago
The paper also fails to show that their central example, Einstein, relied on sensory experience for his intuition leaps rather than general reasoning. They just kind of claim that thought experiments require sensory experience. But you can
16.
▲
by
nearbuy
2mo ago
Only a tiny, tiny fraction of the parameters are encoding information that's specific to a particular programming language. Even if you could remove those without degrading performance, it would have a negligible effect on the model si
17.
▲
by
nearbuy
2mo ago
The problem is Claude Fable is now better than most programmers I know at software architecture and performance optimization as well.
18.
▲
by
nearbuy
2mo ago
That's a good reason to use a dishwasher, and you should keep doing it. But the overall waste is small and people aren't going to care. I'm not saying people who already have a dishwasher will throw it away. But a robot makes
19.
▲
by
nearbuy
2mo ago
Sure, but then there's no such thing as a network that isn't a classifier. Every physically computable function that terminates in finite time will map an input to a fixed set of outputs. And it goes against the common usage, wher
20.
▲
by
nearbuy
2mo ago
The LLM has processed two data modalities derived from the apple (text and vision). Your brain processed a third (taste). But it is still just a data stream, sensing compounds and chemical properties of the apple and turning it into a strea
21.
▲
by
nearbuy
2mo ago
LLMs are not classifiers. A classifier is an algorithm or neural net that assigns a label from a fixed set of labels to an input. You can broaden the definition of classifier to anything that internally divides its input space into regions,
22.
▲
by
nearbuy
2mo ago
I have a dishwasher and I still usually just hand wash. It takes about 10 seconds to wash a dish. The side benefit is all your dishes are always available. With the dishwasher, up to one full dishwasher load are dirty at any time, which mea
23.
▲
by
nearbuy
2mo ago
Unitree's R1 humanoid robot is only about $6000, and it's still a nascent, smallish scale technology. They will come down. If the future home robots are any good, it saves you from buying a dishwasher and robot vacuum. It can repl
24.
▲
by
nearbuy
2mo ago
Yes, much like that. If he hadn't destroyed evidence, he could have argued it was malicious prosecution.
25.
▲
by
nearbuy
2mo ago
What you're describing is malicious prosecution or abuse of process. It's illegal and it would destroy the prosecution's case. Not only that, but the victim could sue for damages.
26.
▲
by
nearbuy
2mo ago
The irony is in this case the in-context and classifier "guardrails" would have almost certainly stopped the attack while their attempts at your definition of guardrails (the sandboxing) failed. In general, people keep trying to m
27.
▲
by
nearbuy
3mo ago
It's a strange experiment. Claude and GPT aren't generating the video. They're directing and editing it, and they request video from a generative video model using mainly text-to-video. Neither Claude nor GPT can actually wat
28.
▲
by
nearbuy
3mo ago
The UNESCO/World Bank literacy rate is basically defined how you thought. But high income countries don't usually report this because literacy by this measure is nearly universal. So they often report at higher thresholds (e.g. ho
29.
▲
by
nearbuy
3mo ago
Questions like that cost a tiny fraction of a cent. "What's the capital of Sri Lanka?" cost a fifth of a cent at GPT 5.5 API price, and would cost a fraction of that if the question were routed to a more suitable, cheaper mod
30.
▲
by
nearbuy
3mo ago
The usage is irrelevant if we're interested in cost per token. If you use it half as much, you get half as many tokens at half the cost. It's still $5.56 in electricity per million output tokens either way (using $0.20/kWh, a
More ›