3 ms·
> It is just an algorithm, we know how it works, We know what calculations it does. We have some hazy idea of some bits of how those calculations lead to somet
by gjm11 13d ago
> It is just an algorithm, we know how it works,
We know what calculations it does. We have some hazy idea of some bits of how those calculations lead to something that at least somewhat resembles intelligent behaviour. But that's a far cry from actually knowing how it works.
For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles (the sort that infamously men are about 1sd better at than women, statistically speaking). How well will it do? I have absolutely no idea and I'm quite sure that a more detailed understanding of the transformer architecture would not make my guesses any better. (Actually, I do kinda have some guesses but they're based on a vague notion about how the models might be partitioned between vision-y bits and language-y bits, and it's very possible that that notion is out of date.)
> it does exactly what we expect it to do
Were you, let's say 6 months ago, expecting it to resolve one of the Millennium Prize problems?
(I do agree that it is more productive to ask "what can and can't they do?" than "should we classify that as intelligent or not?".)
> for Navier-Stokes they spent in 3 days more money than the whole mathematical community over the last 20 years easily.
Are you sure?
(The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.)
- robotpepi 13d ago> Were you, let's say 6 months ago, expecting it to resolve one of the Millennium Prize problems? I didn't expect them to throw millions of dollars at each famous math problem. But one year ago we already had LLMs that solved IMO problems, no? > Are you sure? (The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.) Math has very little founding compared to other science domains. Also, if you filter mathematicians by specialization in PDE and that have worked on Navier-Stokes, then you end up with a very niche community. > For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles. How well will it do? I feel like this is not the correct way of thinking about it. We can also ask, for instance, how well a state-of-the-art algorithm for the salesman problem works on a particular graph topology. People do PhD thesis on topics like that, so the answer is not obvious at all. For LLMs we still don't have a curated theory that explains what they're good/bad at, and that you don't see how to extract an answer from the definitions is no surprise since this is obviously not an easy problem. But all this is normal because this is a rather new topic (models of this scale appeared when? 3 years ago? That's nothing for science). Anthropomorphizing LLMs has added so much noise to this discussion.