Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andy12_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
andy12_
4mo ago
You don't get it. A human set up a software system allowing spicy autocomplete to solve open math problems if the appropriate keyword appears in its output.
32.
▲
by
andy12_
4mo ago
I skimmed through the paper completely expecting polite prompts to do better, and when I saw table 2 I lost it hahahahaha. The rude prompts are specially funny. I mean: > You poor creature, do you even know how to solve this? > Hey go
33.
▲
by
andy12_
5mo ago
Someone blatantly copied their tutorials but ChatGPT is to blame, somehow? The accusation here isn't even that ChatGPT learned from their tutorials and then generated them verbatim. The accusation is that someone copied the whole artic
34.
▲
by
andy12_
5mo ago
> Was the question asked by a mathematician? As per the report, the prompt used to solve the problem is AI-written and the solution was initially graded by an AI grading pipeline. They don't say this explicitly, but it seems like Op
35.
▲
by
andy12_
5mo ago
I disagree. Even frontier models still achieve way worse results than the human baseline in VendingBench. As long as models can't manage optimally something as simple as a vending machine, they have no hope of managing a McDonalds.
36.
▲
by
andy12_
5mo ago
To make performant code sometimes requires implementing or using "unsafe" functions (it's not obligatory, and a lot of projects don't use them; but it was probably needed to map Bun's behavior 1 to 1). Those require
37.
▲
by
andy12_
5mo ago
For now it appears that it talks only to the Codex App. Some users in this thread are saying that apparently the Codex CLI will support it on the next official release.
38.
▲
by
andy12_
5mo ago
Not if you use Linux; app not available yet.
39.
▲
by
andy12_
5mo ago
In my case, because ML research is mainly done with Python+Torch, and if you want people to use your code, you must provide them with python. If it wasn't for that, my dream would be to do ML research in a statically compiled language
40.
▲
by
andy12_
5mo ago
Isn't this already possible to implement with skills and subagents? Like have a skill saying "to test these files run this script that executes a subagent for every markdown file, then check the results".
41.
▲
by
andy12_
5mo ago
It's interesting that Figure 4 shows that Sonnet and Opus have a very clear distinct curve from all other models, even from GPT 5.4. Anthropic superiority I guess.
42.
▲
by
andy12_
5mo ago
>be me >AI goblin-maximizer supervisor >in charge of making sure the AI is, in fact, goblin-maximizing >occasionally have to go down there and check if the AI is still goblin-maximizing >one day i go down there and the AI is
43.
▲
by
andy12_
6mo ago
If you mean for Anthropic in particular, I don't think so. But it's not the first time a major AI lab publishes an incremental update of a model that is worse at some benchmarks. I remember that a particular update of Gemini 2.5 P
44.
▲
by
andy12_
6mo ago
Apparently the score would be a little higher if it weren't for the fact that scores are penalized for being worse than the human baseline, but aren't rewarded for being better than the human baseline (which seems like an arbitrar
45.
▲
by
andy12_
7mo ago
I think that any logic-based test that your average human can "fail" (aka, score below 50%) is not exactly testing for whether something is AGI or not. Though I suppose it depends on your definition of AGI (and whether all humans,
46.
▲
by
andy12_
7mo ago
I think the main value lies in allowing the agent to try many things while you aren't working (when you are sleeping or doing other activities), so even if many tests are not useful, with many trials it can find something nice without
47.
▲
by
andy12_
7mo ago
I think what they mean by this is that, for example, in "If it's raining the outside is wet. It's raining, so the outside is wet", it's more important for the model to learn "If A then B. A, therefore B" t
48.
▲
by
andy12_
7mo ago
Honestly, the most interesting thing here is definitely that just 2D heads are enough to do useful computation (at least they are enough to simulate an interpreter) and that there is an O(log n) algorithm to compute argmax attention with 2D
49.
▲
by
andy12_
7mo ago
This seems a really interesting path for interpretability, specially if a big chunk of a model's behavior occurs pseudo-symbolically. This is an idea I had thought about, integrating tools into the main computation path of a model, but
50.
▲
by
andy12_
7mo ago
There is some things that just don't transfer really well without specific training. I tried to create diagrams in Typst with Cetz (a Processing and Tikz inspired graphing library), and even with documentation, GPT 5.2-thinking can
51.
▲
by
andy12_
7mo ago
> Reality is that we need some way to encode the rules of the world in a more definitive way I mean, sure. But do world models the way LeCun proposes them solves this? I don't think so. JEPAs are just an unsupervised machine learnin
52.
▲
by
andy12_
7mo ago
Putting stuff you have learned into a markdown file is a very "shallow" version of continual learning. It can remember facts, yes, but I doubt a model can master new out-of-distribution tasks this way. If anything, I think that Go
53.
▲
by
andy12_
7mo ago
So, I have been thinking about this for a little while. Image a model f that takes a world x and makes a prediciton y. At a high-level, a traditional supervised model is trained like this f(x)=y' => loss(y',y) => how good wa
54.
▲
by
andy12_
7mo ago
> Even with continuous backpropagation and "learning" That's what I said. Backpropagation cannot be enough; that's not how neurons work in the slightest. When you put biological neurons in a Pong environment they lear
55.
▲
by
andy12_
7mo ago
That's true. Though could that hippocampus-less Einstein be able to keep making novel complex discoveries from that point forward? Seems difficult. He would rapidly reach the limits of his short term memory (the same way current models
56.
▲
by
andy12_
7mo ago
I don't understand this view. How I see it the fundamental bottleneck to AGI is continual learning and backpropagation. Models today are static, and human brains don't learn or adapt themselves with anything close to backpropagati
57.
▲
by
andy12_
7mo ago
Honestly, given that that GPL model would be far below SOTA in capabilities, what exactly would be its use-case? Why would anyone try to use an inferior LLM if they can get away with using a superior one?
58.
▲
by
andy12_
7mo ago
It's not a rumor; you can just test it. Ask the router "What model are you". It will yap on and on about being a GPT-5.3 model (Non-thinking models of OpenAI are insufferable yappers that don't know when to shut up). Ask
59.
▲
by
andy12_
7mo ago
I wasn't prepared to see how good the game looks visually. It's super cool.
60.
▲
by
andy12_
7mo ago
The model could report the confidence of its output distribution, but it isn't necessarily calibrated (that is, even if it tells you that it's 70% confident, it doesn't mean that it is right 70% of the time). Famously, pre-tr
More ›